01 · The Question
If the Dataset Has No Names, How Could Anyone Know Who the Participants Are?
Researchers often remove names before analysis or data sharing. That is sensible, but it can create a false sense of certainty: no names, therefore no identities.
Consider a record describing someone as the 62-year-old female president of the only university in a small municipality. The dataset never states her name. Someone familiar with the institution may not need it.
Names are only one route to identification. Research information can identify people through distinctive characteristics, combinations of variables, codes, contextual clues, metadata, or connections with information held elsewhere.
03 · What You Need to Know
Identification Is About Distinguishing a Person, Not Merely Knowing Their Name
Knowing a Name Is Only One Form of Identification
Identification is broader than attaching a civil name to a record. The Information Commissioner's Office explains that a person may be identifiable even when their name is unknown if they can be distinguished from other people. Its guidance uses singling out and linkability as key indicators when assessing identifiability.
Singling out means being able to isolate the records relating to a particular person or distinguish that person from others. Linkability concerns whether information about the same person can be connected across records or datasets.
This distinction matters because a dataset may consistently distinguish "Participant X" without revealing that person's name. If another source later connects Participant X to a real-world identity, the pieces can come together.
One Distinctive Characteristic May Be Enough
Sometimes identification requires little detective work. A dataset may contain an unusual occupation, public position, rare condition, distinctive award, exceptional age, unique family circumstance, or other characteristic that points to one person in the relevant population.
The relevant population matters. "University president" describes many people nationally. "President of the only university in Municipality X" may describe exactly one.
This is why identifiable research information cannot be defined simply as a fixed list of obvious identifiers.
Several Ordinary Details Can Become Identifying Together
More often, no single variable identifies anyone. Identification emerges from the combination.
Imagine a record containing:
- age: 44;
- occupation: school principal;
- municipality: a small town;
- sex: male; and
- highest degree: doctorate.
Each variable may describe many people when considered independently. Their intersection may describe only one or a few.
This is sometimes called indirect or combinatorial identification. The ICO gives combinations such as age, occupation, and place of residence as examples of criteria that may allow indirect identification when combined with other information.
Consequently, researchers should not assess identifiability one column at a time. The relevant question is what someone could learn from the dataset as a whole.
Public Information Can Supply the Missing Piece
A research file does not need to contain everything necessary for identification. A recipient may already possess the missing information.
University directories identify faculty positions. Professional networking profiles reveal employment histories. Institutional websites announce appointments and awards. Government databases may contain licences or business information. News reports describe public events. Other datasets may contain overlapping attributes.
If the research record and an outside source share sufficiently distinctive information, a person may be identified through linkage.
This is why identification risk can increase substantially when datasets are combined. What appears non-identifying in isolation may become identifying when placed beside another source.
People Who Know the Setting May Need Less Information
Researchers sometimes assess disclosure risk from the perspective of a stranger reading the dataset. That can miss an important audience: people who already know the participants.
Imagine a published quotation:
"After our laboratory was destroyed by the typhoon, I became department chair the following semester."
An outside reader may have no idea who said it. Colleagues in a small department might recognise the person immediately.
This type of contextual identification is especially important in qualitative research, case studies, small communities, workplace research, elite interviews, and studies involving uncommon experiences.
Watch Out
When assessing whether information is identifying, do not ask only whether a stranger could identify the participant. Depending on the applicable framework and disclosure context, colleagues, family members, community members, employers, or other people with relevant background knowledge may have a much easier identification task.
Free Text Can Reintroduce Information You Deliberately Avoided Collecting
A survey may carefully omit names, addresses, and identification numbers, then ask, "Please explain your experience."
Participants can identify themselves or others in the answer. They may mention an employer, supervisor, rare medical condition, specific incident, school, exact date, public achievement, court case, or family relationship.
HHS guidance on HIPAA de-identification specifically notes that identifiers can occur in structured fields or free text and that information-rich clinical narratives may provide context allowing identification. That HIPAA standard applies to a particular US health-information context, but the practical lesson travels well: identifiers do not politely remain in columns labelled "Identifier."
Codes Can Preserve an Identity Connection Without Showing a Name
A dataset may replace names with participant numbers such as P001 and P002. If another file maps those numbers back to identities, the records remain linkable by anyone who can use that file.
This is generally a form of pseudonymization or coding rather than anonymity. Separating the code key from the research data can be an important safeguard, but the existence of the key means an identity route remains.
Researchers therefore need to distinguish pseudonymization from anonymization rather than treating any replacement of names as equivalent to removing identifiability.
Metadata Can Identify People Without Appearing in the Analysis Table
Identification can also arise from information surrounding the dataset. Survey systems may retain IP addresses, account information, timestamps, unique invitation links, or device-related information. Digital files may contain author information or other metadata. Images can contain geolocation or environmental clues.
This matters because researchers often inspect the exported spreadsheet and conclude that it contains no identifiers while overlooking information retained elsewhere in the collection system.
An identifiability assessment should therefore follow the complete data flow, including raw records, platform data, linkage files, metadata, attachments, and exported datasets.
Singling Someone Out Can Matter Even Before Their Name Is Discovered
Suppose a dataset contains a unique pattern that lets someone isolate the same person's records across several years. The observer still does not know the participant's name.
That does not necessarily make the records non-identifiable. The ICO treats singling out as a key indicator of identifiability because being able to distinguish one person can enable further information to be learned about them or later linked to their identity.
This is an important conceptual shift: the question is not only "Can you tell me this person's name?" It is also "Can you distinguish this person from everyone else, connect information about them, or realistically discover who they are?"
Small Populations Make Apparently Ordinary Details More Revealing
The identifying power of information changes with the size and composition of the population.
Knowing that someone is a 35-year-old teacher provides little identifying information in a national dataset. Knowing that someone is the only 35-year-old teacher in a study involving five employees at one rural school is different.
The problem becomes especially acute when researchers report detailed cross-tabulations, individual quotations, rare categories, or combinations of demographics from small research samples.
| Information in the Dataset |
Name Present? |
Possible Identification Route |
| Personal email address |
No |
The address itself may identify the person |
| Age + occupation + small municipality |
No |
The combination may correspond to one person |
| Participant code + accessible code key |
No |
The key restores the identity link |
| Distinctive interview quotation |
No |
People familiar with the participant may recognise the event or wording |
| Precise date + rare event |
No |
External records or news reports may reveal who was involved |
| Unique survey link or account identifier |
No |
System records may connect the response to a participant |
07 · A Quick Checklist
Check Whether Participants Can Be Identified Without Names
Before treating nameless data as non-identifiable, check:
Look for unique or rare occupations, roles, conditions, events, awards, relationships, or other distinctive characteristics.
Test combinations of age, location, occupation, sex or gender, dates, and other demographic variables for uniqueness.
Consider whether someone familiar with the participants or research setting could recognise individuals from contextual clues.
Review free-text responses and quotations for information participants supplied voluntarily.
Check whether participant codes can be linked to identities through a key or another system.
Inspect metadata, timestamps, account information, unique links, IP addresses, filenames, and other system-generated information where relevant.
Consider realistic external sources that could be linked to the research information.
Assess whether individuals can be singled out or their records linked even if their names remain unknown.