Sensitive Data Protection

How to Identify and Classify Sensitive Data with OneTrust

Use OneTrust Data Discovery classification profiles, system and custom classifiers, scans, and review workflows to find and tag sensitive data across your stack.

Working summary: You cannot limit CCPA SPI, apply GDPR Article 9 conditions, or size DPDP safeguards if you do not know where high-harm data sits. OneTrust Data Discovery scans connected sources, classifies findings, and feeds inventories, policies, and rights workflows.

  • Plan labels first: Map classifiers to CCPA SPI, GDPR Article 9, and DPDP high-harm tags before you configure profiles.
  • Profiles: Start from Default or framework packs, clone, and keep only detectors you will review.
  • Custom classifiers: Add regex and dictionaries for internal IDs and product-specific health codes.
  • Review: Use masking and context; accept true positives and escalate ambiguous health or biometric hits.
  • Govern: Map labels to data elements, rights identifiers, encryption tickets, and CCPA limit suppression.

Why classification comes before controls

You cannot limit CCPA sensitive personal information, apply GDPR Article 9 conditions, or size DPDP security safeguards if you do not know where high-harm data sits. OneTrust Data Discovery is built to scan connected sources, classify findings with system and custom classifiers, and feed inventories, policies, and rights workflows.

This article walks through a practical identify-and-classify path using public OneTrust documentation. Product labels and screens change by release; always check your tenant’s MyOneTrust articles for your version. This is not a substitute for OneTrust support or legal advice.

Pair this with the legal maps in GDPR special category data, DPDP Act sensitive data protection, and our CCPA sensitive personal information checklist.

What OneTrust Data Discovery is for

According to OneTrust’s Data Discovery overview, the product scans and classifies data so security and privacy teams can reduce risk from data sprawl, inventory sensitive data, and support governance actions. Public materials describe connectors across structured and unstructured stores, cloud and SaaS systems, and related sources, plus AI-assisted classification and activity mapping.

In practice you will:

  • Connect data sources with appropriate credentials and scan scopes.
  • Attach a classification profile (a bundle of classifiers).
  • Run or schedule scans.
  • Review findings (often with masking and context controls).
  • Map labels to data elements, ROPA or inventory records, and rights-automation identifiers.

Build a short mapping table before you configure classifiers:

  • CCPA / CPRA sensitive personal information (Civil Code section 1798.140): government IDs, credentials with access codes, precise geolocation, certain demographic and belief data, message contents (when you are not the recipient), genetic and neural data, biometrics for unique ID, health, sex life / sexual orientation.
  • GDPR Article 9 special categories: racial or ethnic origin, political opinions, religious or philosophical beliefs, trade union membership, genetic data, biometric data for unique identification, health, sex life / sexual orientation.
  • DPDP: no statutory sensitive tier; still flag high-harm personal data for Rule 6-style safeguards and section 10 sensitivity factors.

One classification run can serve several frameworks if you tag findings with both the OneTrust classifier name and your internal legal tag.

Step 2: Choose or create a classification profile

OneTrust documents “Classification Profiles” as groups of classifiers you run together for a scan. Out-of-the-box profiles described in MyOneTrust include a Default profile aimed at PII generally, a Developer Security profile for keys and tokens, and framework- or region-oriented profiles (for example GDPR-oriented packs with country-specific address patterns).

When you create or edit a profile, documentation describes settings such as:

  • Which system classifiers and custom classifier groups to include.
  • Document classification preferences for unstructured files.
  • Per-classifier sample size and match threshold (structured data; default threshold commonly cited as 30%).
  • Whether a matched label is an identifier for data subject rights automation indexing.

Start with a Default or framework profile, then clone it. Add only the classifiers you will actually review. Too many noisy detectors slow scans and bury real SPI.

Step 3: Understand classifiers and discovery patterns

MyOneTrust “Managing Classifiers” explains that classifiers break broad categories into specific detectors (email, address, national ID, and so on) that you can map back to data elements. Discovery patterns called out in public docs include:

  • Mask and length checks (for example digit patterns for national IDs).
  • Data type and date checks.
  • Lookup lists (dictionaries of known values or terms).
  • Regular expressions.
  • Digital checks to cut false positives.
  • NER / Athena AI patterns when rule-based detectors miss context-heavy text.

Developer documentation for AI Guard also notes large libraries of system classifiers covering personal identifiers, contact data, financial data, healthcare identifiers, credentials, and location patterns. Use that list as a checklist against your CCPA SPI and GDPR Article 9 inventories, not as a legal definition.

Step 4: Add custom classifiers for your business

OneTrust’s custom classifier docs describe building detectors with regex, dictionaries, metadata conditions, and copied templates from system classifiers (system classifiers themselves are not edited in place). Use custom classifiers for:

  • Internal employee or customer ID formats.
  • Country-specific ID patterns not covered by your profile.
  • Product-specific health or loyalty codes that still map to SPI or special category tags in your matrix.

Test patterns on known samples before attaching them to production scans.

Step 5: Connect sources and run scans

  1. Register each database, object store, SaaS app, or file share with least-privilege credentials.
  2. Assign the classification profile to the source or scan job.
  3. Run a limited pilot scan on a non-production copy or a narrow schema first.
  4. Schedule recurring scans for high-churn stores (CRM, support tickets, data lake landing zones).

Document scan IDs and profile versions. MyOneTrust notes you can open a scan job and review which profile and classifiers ran.

Step 6: Review findings with masking and context

Discovery Review workflows let analysts validate terms across sources. OneTrust’s context and masking documentation describes:

  • Masking samples (letters shown as A, digits as N) so reviewers see shape without full SPI.
  • Optional nearby context characters to confirm a match, with a warning that context can itself expose sensitive text.
  • Settings applied on the worker node; existing scan samples do not change until you rescan.

Establish a review queue: accept true positives, reject false positives, and escalate ambiguous health or biometric hits to privacy counsel.

Step 7: Map labels into governance and rights

  • Map classifier labels to your data element catalogue (for example “US SSN” to CCPA SPI government ID and to a security control baseline).
  • Mark identifier classifiers so rights automation can find the same person across systems (OneTrust documents an “Is this an identifier?” style setting on classifiers in a profile).
  • Push inventory updates into records of processing, vendor reviews, and retention jobs.
  • Feed confirmed SPI locations into CCPA limit-the-use suppression lists and into encryption / access-control tickets.

A sensible classifier pack for sensitive data programmes

Use this as a starting checklist, then trim to your stack:

  • National IDs, passport numbers, driver’s licence patterns for markets you serve.
  • Payment card and bank account detectors, plus credential / password / API key packs (Developer Security profile).
  • Email and phone (ordinary PI, still needed for identity resolution).
  • Geolocation / GPS related detectors where your apps collect precise location.
  • Health-related medical record or health plan identifiers if you are a covered health workflow.
  • Custom dictionaries for internal case types that encode diagnosis or orientation.

Do not rely on classifiers alone for political opinions or philosophical beliefs stored as free text; those often need manual ROPA entries and form-design controls.

Operating tips that save rework

  • Separate “detect” from “declare.” A classifier hit is evidence, not a final legal category.
  • Keep EU special category and California SPI tags distinct even when the same column matches both.
  • Rescan after major schema changes, new SaaS tools, or marketing pixel additions.
  • Align scan results with cookie and consent tooling so analytics events do not reintroduce SPI into ad platforms.
  • Version your classification profiles the same way you version infrastructure as code.

Sources

  • OneTrust MyOneTrust: Data Discovery Overview (product purpose and discovery / classification capabilities).
  • OneTrust MyOneTrust: Creating and Managing Classification Profiles (out-of-the-box profiles, thresholds, identifier flag).
  • OneTrust MyOneTrust: Managing Classifiers; Using Custom Classifiers for Discovery; Supported Classifiers and Regions; Context and Masking for Discovery Classification.
  • OneTrust Developer: AI Guard classification profiles (system classifier categories and profile behaviour).
  • California Civil Code section 1798.140 and GDPR Article 9 for legal category maps used in Step 1 (not OneTrust product definitions).

Disclaimer

This article was prepared by Imran using publicly available information. It is for general education only and is not legal advice. Do not rely on it alone when implementing cookie compliance, DSAR handling, consent flows, sensitive-data controls, breach response, or other privacy and data-protection measures. Consult your own legal counsel for advice that fits your business, jurisdictions, and systems.

Last updated on 10 September 2026

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *