01 · The Question
If Identifiers Are Removed, What Should the Data Be Called?
A researcher removes participant names and replaces them with codes. Are the data now de-identified, anonymized, or pseudonymized?
The answer is more complicated than choosing the most technical-sounding term. These concepts overlap, but they do not always describe the same data state. More importantly, their definitions are not perfectly standardized across jurisdictions, disciplines, regulations, and institutional policies.
That variation matters. Calling coded data "anonymous," for example, can substantially misrepresent what has happened if a key still allows the records to be connected to participants.
03 · What You Need to Know
Why These Three Terms Cannot Be Used Interchangeably
Pseudonymization Preserves a Route Back to Identity
Pseudonymization reduces the direct association between research records and participant identities without necessarily eliminating that association permanently.
Under the UK GDPR framework, pseudonymization involves processing personal data so that the data can no longer be attributed to a specific person without additional information, provided that the additional information is kept separately and protected by technical and organisational measures.
A familiar research example is participant coding. Instead of recording "Maria Santos" in the analysis dataset, the researcher records "P047." A separate file states that P047 corresponds to Maria Santos.
Someone with only the research dataset may not know who P047 is. Someone with authorised access to the linking file can restore the connection. The data have therefore not necessarily become anonymous.
Pseudonymized
The identity link has been separated or obscured, but additional information can restore attribution.
Anonymous
The person is no longer identifiable under the applicable standard and circumstances.
The Information Commissioner's Office explicitly cautions against confusing the two. Pseudonymization can reduce risk and improve security, but pseudonymized personal data remain within the scope of UK data-protection law.
Anonymization Aims to Remove Identifiability
Anonymization goes further. Its purpose is to transform information so that it no longer relates to an identified or identifiable person under the applicable standard.
This requires more than replacing names. A participant may remain identifiable through direct or indirect identifiers , combinations of characteristics, contextual knowledge, or linkage with other information.
For that reason, deleting a name while retaining exact birth date, occupation, small geographic area, distinctive experiences, or other identifying details may not achieve effective anonymization.
Anonymization may involve techniques such as suppression, generalisation, aggregation, randomisation, or other statistical disclosure-control methods, depending on the data and intended use. The appropriate approach cannot be reduced to one universal checklist because identification risk depends on the dataset and its environment.
De-Identification Is the Trickiest Term Because Its Meaning Varies
"De-identification" often functions as a broader term for processes that remove, obscure, transform, or otherwise reduce identifying information. NIST, for example, describes de-identification as removing identifying information from a dataset so individual data cannot be linked with specific individuals, while also recognizing that some de-identified data can sometimes be re-identified.
But the term also has specific regulatory meanings. Under the US HIPAA Privacy Rule, protected health information can be designated de-identified through either the Expert Determination method or the Safe Harbor method. Information satisfying the HIPAA de-identification standard is no longer considered protected health information under that rule.
HIPAA's Expert Determination method requires an appropriately qualified person to determine and document that the risk is very small that an anticipated recipient could identify an individual using the information alone or in combination with other reasonably available information. Safe Harbor instead requires removal of specified identifiers and the absence of actual knowledge that the remaining information could identify the person.
Watch Out
Do not treat "de-identified" as a universal technical status with one definition. A paper, ethics application, institution, repository, law, or research team may use the term differently. Identify the governing framework and state what was actually done to the data.
The Terms Do Not Form a Universal Three-Step Ladder
It is tempting to imagine a simple progression:
Identifiable → pseudonymized → de-identified → anonymized.
That sequence may be useful in some local workflows, but it is not a universal taxonomy. Under one framework, de-identification may refer broadly to risk-reduction processes. Under another, it may describe information that has met a particular regulatory threshold. Some organisations also use "de-identified" informally for coded or stripped datasets that would still count as personal data under another legal framework.
Researchers should therefore avoid inferring the status of data from terminology alone. Ask what transformations occurred, whether additional information exists, who possesses it, whether participants remain identifiable, and which definition is being applied.
Coding Is Usually Closer to Pseudonymization Than Anonymization
Consider the common practice of assigning each participant a study ID and storing the identity key separately. This can be an excellent safeguard. It reduces unnecessary exposure of direct identifiers during analysis and can allow researchers to control who has access to identifiable information .
But the key matters. If authorised personnel can use it to determine who P047 is, the connection has been controlled rather than eliminated.
US human-subjects guidance provides a useful illustration. The Office for Human Research Protections describes coded information as information in which identifiers have been replaced by a code while a key exists that permits linkage. Whether investigators themselves can readily ascertain identities can then affect whether information is considered individually identifiable for particular Common Rule purposes.
This also illustrates why identifiability can depend on who holds the data and who can obtain the key.
Removing the Key Does Not Automatically Solve Every Identification Problem
Suppose a research team permanently destroys the code key. That eliminates one route to identification, but the remaining dataset still needs to be assessed.
A record might describe a 74-year-old university president in a particular municipality who received a rare treatment on a precise date. Even without the key, the combination itself might reveal the person.
Anonymization therefore requires researchers to consider identification without explicit names . Destroying the key can materially reduce risk, but the residual information determines whether participants remain identifiable.
Reversibility Is a Useful Practical Question
When terminology becomes confusing, one practical question helps: what would someone need to reconnect this record to a real person?
Term
What Has Happened?
Can Identity Still Be Recovered?
Important Qualification
Pseudonymization
Identifying information is replaced, removed, or separated from other data
Yes, using separately held additional information
The data generally remain personal data under GDPR-style frameworks
Anonymization
Information is transformed so individuals are no longer identifiable under the applicable standard
Identification should not be reasonably achievable under that standard and context
Deleting direct identifiers alone may be insufficient
De-identification
Identifying information or identification risk has been reduced according to a particular process or standard
Depends on the framework and method
The term has no single universal meaning across research contexts
Different Frameworks Can Classify Similar Data Differently
This is not merely a vocabulary dispute. Legal consequences can differ.
Under UK data-protection law, pseudonymized data remain personal data. Under HIPAA, health information meeting the rule's de-identification requirements is no longer protected health information under the Privacy Rule, although HHS explicitly notes that properly de-identified data still retain a very small, non-zero possibility of identification.
Researchers working across countries, institutions, repositories, or collaborative projects should therefore resist importing terminology from one framework into another without checking the definition.
06 · What This Means for You
Describe the Transformation Before Choosing the Label
When preparing a protocol, ethics application, data-management plan, manuscript, or repository documentation, describe what actually happens to the identifying information. The terminology should follow the process, not substitute for explaining it.
A simple decision framework
If names are replaced with codes and a linking key still exists
Pseudonymized or coded data may be the more precise description, subject to the terminology used by your governing framework.
If identifiers have been removed or transformed
Do not infer anonymity automatically. Assess whether participants remain identifiable from residual information or realistic linkage.
If you use the term "de-identified"
Specify the standard or process you mean, especially when a regulation such as HIPAA defines the term.
If no reasonable route to identification remains under the applicable standard
Anonymized or anonymous may be appropriate, provided the conclusion is supported by the required assessment.
Where possible, tell readers whether a key exists, who holds it, what identifiers were transformed, whether indirect identifiers remain, and what release environment is planned. "We pseudonymized the dataset by replacing participant identifiers with random codes and storing the linking key separately" communicates far more than "the data were de-identified."
07 · A Quick Checklist
Before Describing Data as De-Identified, Anonymized, or Pseudonymized
Check the actual data state:
Identify the legal, regulatory, institutional, or technical definition governing the terminology you use.
Document which direct identifiers were removed, replaced, transformed, or separated.
Determine whether a key, codebook, lookup table, or other information can restore participant identities.
Assess whether indirect identifiers and combinations of variables still permit identification.
Check who possesses any additional information needed for re-identification and who can realistically obtain it.
Do not describe pseudonymized data as anonymous merely because analysts cannot see participant names.
For HIPAA-regulated health information, verify whether the relevant Expert Determination or Safe Harbor requirements are actually satisfied.
Describe the transformation and residual identifiability clearly in protocols, consent materials, manuscripts, and data-sharing documentation.
11 · Cite this Guide
How to Cite This Guide
This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.
Recommended (Field Guide)
APA
MLA
Chicago
Copy Citation