Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Are Direct and Indirect Identifiers in Research Data?

Direct identifiers can identify participants relatively directly, while indirect identifiers become identifying through context, combination, or linkage. The distinction helps researchers decide what information requires transformation, separation, or additional protection.

277
Direct and Indirect Identifiers Guide 277 of 398
01 · The Question

Which Research Variables Are Direct Identifiers and Which Are Indirect?

A participant's name is clearly identifying. Age seems less obvious. What about occupation, postcode, exact interview date, photograph, student number, or a rare diagnosis?

Researchers often divide identifying information into direct and indirect identifiers to make this problem more manageable. The distinction is useful, but it is not as simple as memorising two universal lists.

Some information points to a person relatively directly. Other information becomes identifying only through context, combination, or linkage. The same variable can also carry very different identification risk in different datasets.

02 · The Short Answer

Direct Identifiers Point to a Person; Indirect Identifiers Narrow Down Who They Are

In Brief

Direct identifiers can identify or readily point to a particular person on their own or with minimal additional information, while indirect identifiers may identify someone only when combined with other variables, contextual knowledge, or external information.

Names, personal contact details, identification numbers, and some biometric or account identifiers commonly function as direct identifiers. Age, occupation, location, dates, demographic characteristics, and uncommon attributes often function as indirect identifiers, but classification and risk depend on context and the applicable framework.

03 · What You Need to Know

The Difference Is About How Information Leads to Identity

Direct Identifiers Provide a Relatively Immediate Route to a Person

A direct identifier identifies or strongly points to a particular individual without requiring a substantial combination of other research variables.

Common examples include:

  • full name;
  • personal email address;
  • telephone number;
  • student or employee number linked to institutional records;
  • government-issued identification number;
  • medical record number;
  • account number;
  • some biometric identifiers;
  • full-face photographs; and
  • other unique identifiers that directly correspond to an individual.

Not every framework labels exactly the same items "direct identifiers." The practical idea is that these variables provide a relatively short route from the data to the person.

Indirect Identifiers Require Context, Combination, or Linkage

An indirect identifier, sometimes called a quasi-identifier in statistical disclosure contexts, may not reveal a person's identity by itself. Instead, it narrows the population or becomes identifying when combined with other information.

Common examples may include:

  • age or date of birth;
  • sex or gender;
  • occupation;
  • job title or organisational role;
  • geographic area;
  • nationality or ethnicity;
  • educational background;
  • marital or household characteristics;
  • dates associated with events;
  • rare diagnoses or experiences; and
  • other demographic or contextual characteristics.

The ICO illustrates indirect identification using combinations of significant criteria such as age, occupation, and place of residence. The important point is the combination: information that does not directly name someone can still distinguish them when connected with other information.

Indirect Does Not Mean Unimportant

The word "indirect" can sound reassuring, as though these variables are inherently safe. They are not.

HHS guidance on HIPAA de-identification illustrates how seemingly ordinary information can contribute strongly to identification. Under HIPAA's Safe Harbor method, specified geographic detail and most elements of dates directly related to an individual are among the information that must be removed, along with names and numerous conventional identifiers. The rule also requires that the covered entity have no actual knowledge that the remaining information could identify an individual alone or in combination with other information.

HIPAA is a specific US health-information framework, so its list should not be copied into every research protocol as a universal taxonomy. It does, however, demonstrate why dates, geography, and unique characteristics deserve attention alongside obvious identifiers.

Several Weak Identifiers Can Become One Strong Identification Pattern

Consider four variables:

  • age: 53;
  • occupation: veterinarian;
  • municipality: a small town; and
  • sex: female.

None necessarily identifies a participant independently. If only one 53-year-old female veterinarian lives or works in that town, their combination may effectively single her out.

This is why researchers should assess whether participants can be identified without names. The identifying unit may be a pattern rather than a field.

Direct and Indirect Are Not Always Permanent Labels

The classification of an identifier can depend on who holds it and what systems are available.

A random-looking employee number may mean nothing to an outside researcher. To the employer with access to a personnel database, the number may immediately reveal a person's identity. A vehicle registration number may similarly identify a vehicle directly while identifying its owner through access to registration information.

The ICO specifically notes that identifiers such as vehicle registration numbers, passport numbers, and combinations of significant criteria can enable indirect identification when connected with additional information.

Researchers should therefore use the direct/indirect distinction as an analytical tool rather than assuming that every variable carries one immutable label in every environment.

Unique Codes Need Special Treatment

Study IDs are deliberately created to replace direct identifiers during research. Whether they should be treated as identifying depends partly on the surrounding system.

If P004 can be connected to a participant using a code key, the code participates in an identity-linkage mechanism. It may be meaningless to an analyst who lacks the key while remaining linkable to the research team that holds it.

Separating identifiers and linkage information can therefore be valuable even though it does not automatically make the underlying data anonymous. Researchers may need to consider when participant identifiers should be stored separately.

Context Determines How Revealing an Indirect Identifier Is

Suppose a dataset records participants' occupations.

"Teacher" may reveal little in a large national survey. "Only pediatric neurosurgeon at Hospital X" is another matter. Likewise, a province may contain millions of residents, while a tiny barangay, postcode, or institutional unit may narrow the population substantially.

The identifying power of an indirect identifier depends on factors such as:

  • how detailed the variable is;
  • how common or rare the value is;
  • the size of the relevant population;
  • which other variables accompany it;
  • what outside information exists; and
  • who receives the data.

This is why demographic information can accidentally identify participants in some studies while remaining relatively non-distinctive in others.

Identifiers Can Hide in Unstructured Data

Direct and indirect identifiers are not confined to tidy spreadsheet columns.

An interview participant may name their employer. A diary may contain an exact address. A clinical narrative may mention an unusual occupation. A photograph may reveal a face or licence plate. An audio recording may contain a recognisable voice. A video may show a distinctive workplace.

HHS guidance explicitly notes that HIPAA identifiers must be addressed whether they appear in structured fields or free text. It gives information-rich clinical narratives as an example of material that can contain contextual information allowing identification.

Qualitative researchers therefore need to inspect informational content, not merely remove columns from a spreadsheet.

Direct and Indirect Identifiers Differ From Sensitive Information

Identifiability and sensitivity are related but separate dimensions.

Identifier Information that helps distinguish or determine who a person is.
Sensitive information Information whose disclosure may create particular privacy, social, legal, economic, or other consequences, depending on context and applicable law.

A person's telephone number may identify them without revealing a sensitive research finding. A highly sensitive opinion may reveal very little about identity if it is genuinely disconnected from the person who expressed it.

When the two occur together, the consequences of identification can increase considerably. Researchers should therefore assess both the likelihood of identification and the consequences if identification occurs.

A Practical Taxonomy Helps, but It Is Not a Universal Legal List

Information Common Research Classification Why It Can Identify
Full name Direct Explicitly names the person
Personal email address Direct Often corresponds to a specific person
Government or institutional ID Direct or readily linkable Can connect directly to administrative records
Full-face photograph Direct or strongly identifying Visual recognition can reveal identity
Age Usually indirect Narrows the possible population
Occupation Usually indirect May become distinctive in a small population
Geographic area Usually indirect, depending on precision Reduces the relevant population and supports linkage
Exact event date Often indirect May connect a record with external events or records
Rare characteristic Indirect but potentially powerful May single out one or very few people
Study code with a retained key Linkable identifier The key can restore attribution to a participant

Use such classifications to structure a risk assessment, not as a substitute for checking the definitions imposed by your ethics body, institution, jurisdiction, or data-sharing framework.

04 · A Practical Example

How Indirect Identifiers Can Reconstruct an Identity

Hypothetical Example

A Survey of School Leaders

A researcher surveys school leaders about burnout. Names, telephone numbers, and email addresses are stored separately from the analysis dataset.

Direct identifiers The contact file contains each participant's name and email address. These provide a straightforward route to identity and are separated from survey responses.
Indirect identifiers The analysis dataset contains age, sex, school type, municipality, position, years in the role, and highest qualification.
The combination One participant is a 39-year-old female superintendent with a doctorate in a small municipality containing only one person in that position.
The result The record may be identifiable from the indirect variables even though the direct identifiers are stored elsewhere.
The response The researcher considers whether every demographic detail is analytically necessary and whether categories can be broadened or particular variables suppressed in shared or published outputs.

The exercise shows why separating direct identifiers is useful but not sufficient for every disclosure problem. The remaining variables still need their own identifiability assessment.

05 · What Researchers Often Get Wrong

Common Mistakes About Direct and Indirect Identifiers

Misconception

Only Direct Identifiers Can Reveal Identity

Indirect identifiers can identify a participant through combination, contextual knowledge, or linkage. Their individual weakness should not be confused with collective harmlessness.

Misconception

Indirect Identifiers Are Always Demographics

No. Dates, geographic information, institutional roles, unusual events, rare characteristics, metadata, behavioural patterns, and other contextual information may also function indirectly.

Misconception

Every Variable Has One Permanent Classification

The same information may have different identifying power depending on the recipient, available databases, population, precision, and research environment. Classification should reflect context rather than vocabulary alone.

Misconception

Removing Direct Identifiers Makes the Dataset Anonymous

It reduces an important class of identification risk, but indirect identifiers and linkage may remain. Anonymity requires a broader assessment of whether individuals are still identifiable.

Misconception

If a Variable Is Useful, It Must Be Safe to Retain at Full Precision

Analytical usefulness and identification risk are separate considerations. Researchers may sometimes preserve necessary information while reducing precision, grouping categories, restricting access, or using another disclosure-control approach appropriate to the study.

06 · What This Means for You

Review Identifiers as a System, Not as Individual Columns

Start by identifying obvious direct identifiers. Then examine what remains collectively. The second stage is where many disclosure risks become visible.

A simple decision framework

If a direct identifier is unnecessary
Do not collect it, or remove it when it no longer serves an approved purpose.
If a direct identifier is necessary only for contact or linkage
Consider storing it separately from substantive research responses and limiting access appropriately.
If an indirect identifier is analytically necessary
Assess its precision, rarity, and interaction with the other variables rather than deleting it automatically.
If several indirect identifiers create distinctive combinations
Consider generalisation, suppression, aggregation, access restrictions, or another suitable disclosure-control strategy.
If the dataset will be shared with a new recipient
Reconsider what additional information that recipient may possess and whether the identifying power of the variables changes.

The practical goal is not to purge every potentially identifying variable regardless of scientific value. It is to collect and retain what the research genuinely requires, understand how those variables interact, and protect them in proportion to the identification risk they create.

07 · A Quick Checklist

Audit Both Direct and Indirect Identifiers

When reviewing a research dataset, check:
Identify names, personal contact information, identification numbers, account information, photographs, and other obvious direct identifiers.
Review age, dates, geography, occupation, institutional role, and demographic characteristics as potential indirect identifiers.
Examine combinations of indirect identifiers rather than assessing each variable independently.
Look for rare values and unique characteristics that could single out one participant.
Check study codes and determine whether a linkage key or other mechanism can restore identities.
Review free text, images, audio, video, and metadata for identifiers outside structured variables.
Consider external information and datasets that realistic recipients could use for linkage.
Verify whether each identifying variable is genuinely necessary at its current level of precision.
Apply the definitions and requirements of the legal, ethical, institutional, or data-sharing framework governing the study.
08 · Frequently Asked Questions

Frequently Asked Questions About Direct and Indirect Identifiers

Is age a direct or indirect identifier?

Age is commonly treated as an indirect identifier because it usually narrows the population rather than naming one person. Very unusual ages or age combined with other information can nevertheless become highly identifying.

Is date of birth a direct identifier?

Classification varies by framework and context. A date of birth may not uniquely identify someone by itself, but it can be a powerful identifier when combined with other information. Some regulatory de-identification standards specifically restrict detailed dates because of their identification risk.

Is an email address a direct identifier?

A personal email address commonly functions as a direct identifier because it is often associated with a particular person. Generic or shared addresses may behave differently, so context still matters.

Is a postcode or ZIP code an indirect identifier?

Geographic information often functions as an indirect identifier because it narrows the relevant population. Identification risk increases as geographic precision increases and when geography is combined with other characteristics.

Is a participant code a direct identifier?

A random study code may reveal nothing by itself, but if a key connects it to a participant, it remains part of an identification mechanism. Its practical status therefore depends on who can access the linking information.

How many indirect identifiers does it take to identify someone?

There is no universal number. Identification depends on the values' rarity, precision, population size, other available information, and the recipient's knowledge. One exceptionally distinctive attribute may sometimes be enough, while several broad attributes may still describe many people.

Should researchers remove every indirect identifier?

No. Some indirect identifiers may be essential to the research. Researchers should assess necessity and identification risk, then consider appropriate transformations, access controls, or other safeguards rather than automatically deleting scientifically useful information.

09 · The Bottom Line

Indirect Identifiers Matter Most in Combination

The Bottom Line

Direct identifiers provide a relatively immediate route to a participant's identity, while indirect identifiers can reveal identity through combination, context, singling out, or linkage with other information.

Remove or separate unnecessary direct identifiers, but do not stop there. Review the remaining variables collectively and in their actual data environment, because a dataset full of individually ordinary details can still describe an extraordinary one-person group.

10 · Sources and Further Reading

Authoritative Sources on Direct and Indirect Identification

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes