Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

De-Identification vs. Anonymization vs. Pseudonymization: What’s the Difference?

De-identification, anonymization, and pseudonymization all reduce exposure of identity, but they are not interchangeable. Their meanings also depend partly on the legal, regulatory, and technical framework in which the terms are used.

274
De-Identification vs. Anonymization vs. Pseudonymization Guide 274 of 398
01 · The Question

If Identifiers Are Removed, What Should the Data Be Called?

A researcher removes participant names and replaces them with codes. Are the data now de-identified, anonymized, or pseudonymized?

The answer is more complicated than choosing the most technical-sounding term. These concepts overlap, but they do not always describe the same data state. More importantly, their definitions are not perfectly standardized across jurisdictions, disciplines, regulations, and institutional policies.

That variation matters. Calling coded data "anonymous," for example, can substantially misrepresent what has happened if a key still allows the records to be connected to participants.

02 · The Short Answer

The Terms Describe Different Ways of Reducing Identifiability

In Brief

Pseudonymization replaces or separates identifying information while preserving a way to reconnect data to individuals; anonymization aims to make individuals no longer identifiable under the applicable standard; and de-identification is a broader term whose precise meaning depends on the framework in which it is used.

Do not assume that "de-identified" always means anonymous. In some frameworks, de-identification has a specific legal standard; elsewhere it may describe a range of techniques that reduce identifiability without necessarily eliminating the possibility of re-identification.

03 · What You Need to Know

Why These Three Terms Cannot Be Used Interchangeably

Pseudonymization Preserves a Route Back to Identity

Pseudonymization reduces the direct association between research records and participant identities without necessarily eliminating that association permanently.

Under the UK GDPR framework, pseudonymization involves processing personal data so that the data can no longer be attributed to a specific person without additional information, provided that the additional information is kept separately and protected by technical and organisational measures.

A familiar research example is participant coding. Instead of recording "Maria Santos" in the analysis dataset, the researcher records "P047." A separate file states that P047 corresponds to Maria Santos.

Someone with only the research dataset may not know who P047 is. Someone with authorised access to the linking file can restore the connection. The data have therefore not necessarily become anonymous.

Pseudonymized The identity link has been separated or obscured, but additional information can restore attribution.
Anonymous The person is no longer identifiable under the applicable standard and circumstances.

The Information Commissioner's Office explicitly cautions against confusing the two. Pseudonymization can reduce risk and improve security, but pseudonymized personal data remain within the scope of UK data-protection law.

Anonymization Aims to Remove Identifiability

Anonymization goes further. Its purpose is to transform information so that it no longer relates to an identified or identifiable person under the applicable standard.

This requires more than replacing names. A participant may remain identifiable through direct or indirect identifiers, combinations of characteristics, contextual knowledge, or linkage with other information.

For that reason, deleting a name while retaining exact birth date, occupation, small geographic area, distinctive experiences, or other identifying details may not achieve effective anonymization.

Anonymization may involve techniques such as suppression, generalisation, aggregation, randomisation, or other statistical disclosure-control methods, depending on the data and intended use. The appropriate approach cannot be reduced to one universal checklist because identification risk depends on the dataset and its environment.

De-Identification Is the Trickiest Term Because Its Meaning Varies

"De-identification" often functions as a broader term for processes that remove, obscure, transform, or otherwise reduce identifying information. NIST, for example, describes de-identification as removing identifying information from a dataset so individual data cannot be linked with specific individuals, while also recognizing that some de-identified data can sometimes be re-identified.

But the term also has specific regulatory meanings. Under the US HIPAA Privacy Rule, protected health information can be designated de-identified through either the Expert Determination method or the Safe Harbor method. Information satisfying the HIPAA de-identification standard is no longer considered protected health information under that rule.

HIPAA's Expert Determination method requires an appropriately qualified person to determine and document that the risk is very small that an anticipated recipient could identify an individual using the information alone or in combination with other reasonably available information. Safe Harbor instead requires removal of specified identifiers and the absence of actual knowledge that the remaining information could identify the person.

Watch Out

Do not treat "de-identified" as a universal technical status with one definition. A paper, ethics application, institution, repository, law, or research team may use the term differently. Identify the governing framework and state what was actually done to the data.

The Terms Do Not Form a Universal Three-Step Ladder

It is tempting to imagine a simple progression:

Identifiable → pseudonymized → de-identified → anonymized.

That sequence may be useful in some local workflows, but it is not a universal taxonomy. Under one framework, de-identification may refer broadly to risk-reduction processes. Under another, it may describe information that has met a particular regulatory threshold. Some organisations also use "de-identified" informally for coded or stripped datasets that would still count as personal data under another legal framework.

Researchers should therefore avoid inferring the status of data from terminology alone. Ask what transformations occurred, whether additional information exists, who possesses it, whether participants remain identifiable, and which definition is being applied.

Coding Is Usually Closer to Pseudonymization Than Anonymization

Consider the common practice of assigning each participant a study ID and storing the identity key separately. This can be an excellent safeguard. It reduces unnecessary exposure of direct identifiers during analysis and can allow researchers to control who has access to identifiable information.

But the key matters. If authorised personnel can use it to determine who P047 is, the connection has been controlled rather than eliminated.

US human-subjects guidance provides a useful illustration. The Office for Human Research Protections describes coded information as information in which identifiers have been replaced by a code while a key exists that permits linkage. Whether investigators themselves can readily ascertain identities can then affect whether information is considered individually identifiable for particular Common Rule purposes.

This also illustrates why identifiability can depend on who holds the data and who can obtain the key.

Removing the Key Does Not Automatically Solve Every Identification Problem

Suppose a research team permanently destroys the code key. That eliminates one route to identification, but the remaining dataset still needs to be assessed.

A record might describe a 74-year-old university president in a particular municipality who received a rare treatment on a precise date. Even without the key, the combination itself might reveal the person.

Anonymization therefore requires researchers to consider identification without explicit names. Destroying the key can materially reduce risk, but the residual information determines whether participants remain identifiable.

Reversibility Is a Useful Practical Question

When terminology becomes confusing, one practical question helps: what would someone need to reconnect this record to a real person?

Term What Has Happened? Can Identity Still Be Recovered? Important Qualification
Pseudonymization Identifying information is replaced, removed, or separated from other data Yes, using separately held additional information The data generally remain personal data under GDPR-style frameworks
Anonymization Information is transformed so individuals are no longer identifiable under the applicable standard Identification should not be reasonably achievable under that standard and context Deleting direct identifiers alone may be insufficient
De-identification Identifying information or identification risk has been reduced according to a particular process or standard Depends on the framework and method The term has no single universal meaning across research contexts

Different Frameworks Can Classify Similar Data Differently

This is not merely a vocabulary dispute. Legal consequences can differ.

Under UK data-protection law, pseudonymized data remain personal data. Under HIPAA, health information meeting the rule's de-identification requirements is no longer protected health information under the Privacy Rule, although HHS explicitly notes that properly de-identified data still retain a very small, non-zero possibility of identification.

Researchers working across countries, institutions, repositories, or collaborative projects should therefore resist importing terminology from one framework into another without checking the definition.

04 · A Practical Example

One Dataset Can Pass Through Several Different States

Hypothetical Example

A Longitudinal Student Well-Being Study

A research team collects names, university email addresses, demographic information, and well-being responses. Participants will complete three surveys over two years, so researchers need to connect each person's responses across waves.

Original identifiable data The intake records directly contain participants' names and contact details alongside their study information.
Pseudonymized working dataset The team assigns each participant a random study code. Names and contact information are moved into a separately protected linking file. Analysts receive only the coded dataset.
After longitudinal follow-up The research team no longer needs the identity link and permanently removes the linking information according to the approved data-management plan.
Anonymization assessment The team does not simply declare the remaining file anonymous. It evaluates whether combinations of demographics, dates, free-text responses, or other information could still identify participants.
Data sharing If the residual identification risk remains too high for open release, the team may further transform the data or use controlled access rather than calling the dataset anonymous prematurely.

The important point is that removing direct identifiers, assigning codes, destroying a key, and achieving effective anonymization are conceptually different actions. A good data-management plan says which one is happening and when.

05 · What Researchers Often Get Wrong

Common Mistakes With De-Identification Terminology

Misconception

De-Identified Always Means Anonymous

Not across all research contexts. "De-identification" can have a specific regulatory definition or be used more broadly for techniques that reduce identifiability. State the governing standard and what was actually done.

Misconception

Replacing Names With Codes Anonymizes the Data

If a key or other information can reconnect codes to participants, the process is ordinarily closer to pseudonymization. Separating the key can substantially improve protection without eliminating the identity link.

Misconception

Pseudonymized Data Are No Longer Personal Data

Under GDPR-style frameworks, pseudonymized personal data remain personal data because individuals can still be attributed using additional information. Pseudonymization is a safeguard, not an automatic route outside data-protection law.

Misconception

Destroying the Code Key Automatically Makes Data Anonymous

Destroying the key removes an important linkage mechanism, but participants may remain identifiable from the residual dataset itself or through other available information. The remaining data still require an identifiability assessment.

Misconception

There Is One International Definition of De-Identification

No single definition governs every research context. HIPAA, GDPR-style regimes, human-subjects regulations, statistical agencies, repositories, and institutions may use different concepts and thresholds. Researchers should verify the framework that actually applies.

06 · What This Means for You

Describe the Transformation Before Choosing the Label

When preparing a protocol, ethics application, data-management plan, manuscript, or repository documentation, describe what actually happens to the identifying information. The terminology should follow the process, not substitute for explaining it.

A simple decision framework

If names are replaced with codes and a linking key still exists
Pseudonymized or coded data may be the more precise description, subject to the terminology used by your governing framework.
If identifiers have been removed or transformed
Do not infer anonymity automatically. Assess whether participants remain identifiable from residual information or realistic linkage.
If you use the term "de-identified"
Specify the standard or process you mean, especially when a regulation such as HIPAA defines the term.
If no reasonable route to identification remains under the applicable standard
Anonymized or anonymous may be appropriate, provided the conclusion is supported by the required assessment.

Where possible, tell readers whether a key exists, who holds it, what identifiers were transformed, whether indirect identifiers remain, and what release environment is planned. "We pseudonymized the dataset by replacing participant identifiers with random codes and storing the linking key separately" communicates far more than "the data were de-identified."

07 · A Quick Checklist

Before Describing Data as De-Identified, Anonymized, or Pseudonymized

Check the actual data state:
Identify the legal, regulatory, institutional, or technical definition governing the terminology you use.
Document which direct identifiers were removed, replaced, transformed, or separated.
Determine whether a key, codebook, lookup table, or other information can restore participant identities.
Assess whether indirect identifiers and combinations of variables still permit identification.
Check who possesses any additional information needed for re-identification and who can realistically obtain it.
Do not describe pseudonymized data as anonymous merely because analysts cannot see participant names.
For HIPAA-regulated health information, verify whether the relevant Expert Determination or Safe Harbor requirements are actually satisfied.
Describe the transformation and residual identifiability clearly in protocols, consent materials, manuscripts, and data-sharing documentation.
08 · Frequently Asked Questions

Frequently Asked Questions About De-Identification and Anonymization

Are pseudonymized data anonymous?

No, not when additional information can still attribute the records to individuals under the applicable framework. Pseudonymization reduces direct identifiability while preserving a controlled route to attribution.

Is coded research data pseudonymized?

Often, yes, when identifiers are replaced with codes and a separate key permits linkage. Terminology varies across frameworks, so researchers should verify the definition used by their institution or applicable regulation.

Is de-identified data always non-identifiable?

Not as a universal proposition because "de-identified" is used differently across frameworks. Under HIPAA it has a specific regulatory standard, while other research contexts may use the term more broadly. State the standard being applied.

Does HIPAA Safe Harbor mean removing only names?

No. HIPAA Safe Harbor specifies multiple categories of identifiers that must be removed and also requires that the covered entity have no actual knowledge that the remaining information could identify the individual.

Can pseudonymization still be useful if the data remain identifiable?

Yes. It can substantially reduce unnecessary exposure of identities, support access controls, limit the consequences of some data breaches, and facilitate research processing while keeping the linking information separately protected.

Can pseudonymized data later become anonymous?

Potentially. Removing the linkage mechanism may be part of anonymization, but researchers must also assess whether the remaining data permit identification through other information or realistic linkage.

09 · The Bottom Line

Do Not Let Similar Terminology Hide Different Data States

The Bottom Line

Pseudonymization preserves a controlled route back to identity, anonymization aims to remove identifiability under the applicable standard, and de-identification must be interpreted according to the specific framework in which the term is used.

When precision matters, describe what happened to the identifiers, whether additional information can restore identity, and which standard you are applying. A clear description of the data is more reliable than assuming everyone means the same thing by "de-identified."

10 · Sources and Further Reading

Authoritative Sources on De-Identification and Pseudonymization

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes