Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Can Someone Be Identified Without Their Name Appearing in the Dataset?

Removing names does not necessarily make research data non-identifiable. Participants may still be singled out or identified through combinations of characteristics, contextual knowledge, metadata, or links to other information.

276
Identification Without Names Guide 276 of 398
01 · The Question

If the Dataset Has No Names, How Could Anyone Know Who the Participants Are?

Researchers often remove names before analysis or data sharing. That is sensible, but it can create a false sense of certainty: no names, therefore no identities.

Consider a record describing someone as the 62-year-old female president of the only university in a small municipality. The dataset never states her name. Someone familiar with the institution may not need it.

Names are only one route to identification. Research information can identify people through distinctive characteristics, combinations of variables, codes, contextual clues, metadata, or connections with information held elsewhere.

02 · The Short Answer

Yes, a Person Can Be Identified Without Being Named

In Brief

A research participant can be identifiable even when their name never appears in the dataset because other information may single them out or allow their record to be connected to their identity.

Identification may arise from one distinctive attribute, a combination of ordinary variables, a participant code, free-text information, metadata, contextual knowledge, or linkage with another dataset. Whether this is realistically possible depends on the information, the population, who has access, and what additional information is available.

03 · What You Need to Know

Identification Is About Distinguishing a Person, Not Merely Knowing Their Name

Knowing a Name Is Only One Form of Identification

Identification is broader than attaching a civil name to a record. The Information Commissioner's Office explains that a person may be identifiable even when their name is unknown if they can be distinguished from other people. Its guidance uses singling out and linkability as key indicators when assessing identifiability.

Singling out means being able to isolate the records relating to a particular person or distinguish that person from others. Linkability concerns whether information about the same person can be connected across records or datasets.

This distinction matters because a dataset may consistently distinguish "Participant X" without revealing that person's name. If another source later connects Participant X to a real-world identity, the pieces can come together.

One Distinctive Characteristic May Be Enough

Sometimes identification requires little detective work. A dataset may contain an unusual occupation, public position, rare condition, distinctive award, exceptional age, unique family circumstance, or other characteristic that points to one person in the relevant population.

The relevant population matters. "University president" describes many people nationally. "President of the only university in Municipality X" may describe exactly one.

This is why identifiable research information cannot be defined simply as a fixed list of obvious identifiers.

Several Ordinary Details Can Become Identifying Together

More often, no single variable identifies anyone. Identification emerges from the combination.

Imagine a record containing:

  • age: 44;
  • occupation: school principal;
  • municipality: a small town;
  • sex: male; and
  • highest degree: doctorate.

Each variable may describe many people when considered independently. Their intersection may describe only one or a few.

This is sometimes called indirect or combinatorial identification. The ICO gives combinations such as age, occupation, and place of residence as examples of criteria that may allow indirect identification when combined with other information.

Consequently, researchers should not assess identifiability one column at a time. The relevant question is what someone could learn from the dataset as a whole.

Public Information Can Supply the Missing Piece

A research file does not need to contain everything necessary for identification. A recipient may already possess the missing information.

University directories identify faculty positions. Professional networking profiles reveal employment histories. Institutional websites announce appointments and awards. Government databases may contain licences or business information. News reports describe public events. Other datasets may contain overlapping attributes.

If the research record and an outside source share sufficiently distinctive information, a person may be identified through linkage.

This is why identification risk can increase substantially when datasets are combined. What appears non-identifying in isolation may become identifying when placed beside another source.

People Who Know the Setting May Need Less Information

Researchers sometimes assess disclosure risk from the perspective of a stranger reading the dataset. That can miss an important audience: people who already know the participants.

Imagine a published quotation:

"After our laboratory was destroyed by the typhoon, I became department chair the following semester."

An outside reader may have no idea who said it. Colleagues in a small department might recognise the person immediately.

This type of contextual identification is especially important in qualitative research, case studies, small communities, workplace research, elite interviews, and studies involving uncommon experiences.

Watch Out

When assessing whether information is identifying, do not ask only whether a stranger could identify the participant. Depending on the applicable framework and disclosure context, colleagues, family members, community members, employers, or other people with relevant background knowledge may have a much easier identification task.

Free Text Can Reintroduce Information You Deliberately Avoided Collecting

A survey may carefully omit names, addresses, and identification numbers, then ask, "Please explain your experience."

Participants can identify themselves or others in the answer. They may mention an employer, supervisor, rare medical condition, specific incident, school, exact date, public achievement, court case, or family relationship.

HHS guidance on HIPAA de-identification specifically notes that identifiers can occur in structured fields or free text and that information-rich clinical narratives may provide context allowing identification. That HIPAA standard applies to a particular US health-information context, but the practical lesson travels well: identifiers do not politely remain in columns labelled "Identifier."

Codes Can Preserve an Identity Connection Without Showing a Name

A dataset may replace names with participant numbers such as P001 and P002. If another file maps those numbers back to identities, the records remain linkable by anyone who can use that file.

This is generally a form of pseudonymization or coding rather than anonymity. Separating the code key from the research data can be an important safeguard, but the existence of the key means an identity route remains.

Researchers therefore need to distinguish pseudonymization from anonymization rather than treating any replacement of names as equivalent to removing identifiability.

Metadata Can Identify People Without Appearing in the Analysis Table

Identification can also arise from information surrounding the dataset. Survey systems may retain IP addresses, account information, timestamps, unique invitation links, or device-related information. Digital files may contain author information or other metadata. Images can contain geolocation or environmental clues.

This matters because researchers often inspect the exported spreadsheet and conclude that it contains no identifiers while overlooking information retained elsewhere in the collection system.

An identifiability assessment should therefore follow the complete data flow, including raw records, platform data, linkage files, metadata, attachments, and exported datasets.

Singling Someone Out Can Matter Even Before Their Name Is Discovered

Suppose a dataset contains a unique pattern that lets someone isolate the same person's records across several years. The observer still does not know the participant's name.

That does not necessarily make the records non-identifiable. The ICO treats singling out as a key indicator of identifiability because being able to distinguish one person can enable further information to be learned about them or later linked to their identity.

This is an important conceptual shift: the question is not only "Can you tell me this person's name?" It is also "Can you distinguish this person from everyone else, connect information about them, or realistically discover who they are?"

Small Populations Make Apparently Ordinary Details More Revealing

The identifying power of information changes with the size and composition of the population.

Knowing that someone is a 35-year-old teacher provides little identifying information in a national dataset. Knowing that someone is the only 35-year-old teacher in a study involving five employees at one rural school is different.

The problem becomes especially acute when researchers report detailed cross-tabulations, individual quotations, rare categories, or combinations of demographics from small research samples.

Information in the Dataset Name Present? Possible Identification Route
Personal email address No The address itself may identify the person
Age + occupation + small municipality No The combination may correspond to one person
Participant code + accessible code key No The key restores the identity link
Distinctive interview quotation No People familiar with the participant may recognise the event or wording
Precise date + rare event No External records or news reports may reveal who was involved
Unique survey link or account identifier No System records may connect the response to a participant
04 · A Practical Example

How Six Innocent-Looking Details Can Reveal One Participant

Hypothetical Example

A Study of Academic Leadership Experiences

A researcher interviews senior academics from several institutions. Names, email addresses, and university names are removed before transcripts are shared with collaborators.

The transcript One participant is described as a 41-year-old woman, dean of a health-related college, appointed in 2025, at a university in a small city.
An additional clue The participant discusses receiving a distinctive national professional award shortly before becoming dean.
The external search A university announcement identifies a dean appointed in 2025, while the professional association's website lists the previous year's award recipient.
The match The demographic, professional, geographic, and award information all point to the same person.
The lesson No name was necessary. Identity emerged from the intersection between the research information and information available elsewhere.

Reducing this risk might involve removing unnecessary detail, broadening age or geography, generalising the appointment period, modifying the description of the award, restricting access to the full transcript, or using another disclosure-control strategy appropriate to the research purpose.

05 · What Researchers Often Get Wrong

Common Misunderstandings About Identification Without Names

Misconception

No Name Means No Identification

A name is only one identifier. A unique characteristic, combination of variables, code, metadata item, quotation, or external linkage may identify someone without their name appearing anywhere in the research file.

Misconception

If I Cannot Identify the Participant, Nobody Else Can

Other people may possess information that the researcher does not. Colleagues, community members, data recipients, or organisations with additional datasets may be able to make connections unavailable to the researcher.

Misconception

Ordinary Demographic Variables Are Harmless

Broad demographic categories may carry little identifying power individually, but combinations can become highly distinctive. Their risk depends on granularity, population size, and what other information is available.

Misconception

If Nobody Knows the Person's Name, the Record Cannot Be Personal

A person can sometimes be singled out, tracked, or distinguished from others without their name being known. Depending on the governing framework, that may still constitute identifiability.

Misconception

Deleting Identifiers From the Spreadsheet Solves the Problem

Identifying information may remain in free text, recordings, metadata, code keys, survey-platform records, filenames, or other materials outside the analysis spreadsheet. Assess the whole research information environment.

06 · What This Means for You

Ask How Someone Could Work Backward From the Data to the Person

When reviewing a dataset, adopt the perspective of a potential recipient rather than checking only whether obvious identifiers were removed. What could that person infer from the remaining information? What might they already know? What external information could realistically complete the identification?

A simple decision framework

If one attribute uniquely describes a participant
Consider whether the level of detail is necessary or whether it can be generalised, suppressed, or otherwise protected.
If no individual variable identifies anyone
Examine combinations of variables before concluding that the dataset is non-identifiable.
If participants belong to a small or distinctive population
Assume contextual knowledge may make apparently ordinary details more revealing and assess disclosure accordingly.
If data will be shared beyond the original team
Consider what additional information recipients may possess and whether the intended access model changes identification risk.
If free text, audio, images, or metadata are present
Review those materials specifically rather than assuming structured-variable de-identification covers them.

The goal is not to imagine an omniscient adversary with unlimited resources. It is to assess plausible identification routes under the standard and circumstances that apply to the study. That produces a more defensible assessment than treating names as an on/off switch for identifiability.

07 · A Quick Checklist

Check Whether Participants Can Be Identified Without Names

Before treating nameless data as non-identifiable, check:
Look for unique or rare occupations, roles, conditions, events, awards, relationships, or other distinctive characteristics.
Test combinations of age, location, occupation, sex or gender, dates, and other demographic variables for uniqueness.
Consider whether someone familiar with the participants or research setting could recognise individuals from contextual clues.
Review free-text responses and quotations for information participants supplied voluntarily.
Check whether participant codes can be linked to identities through a key or another system.
Inspect metadata, timestamps, account information, unique links, IP addresses, filenames, and other system-generated information where relevant.
Consider realistic external sources that could be linked to the research information.
Assess whether individuals can be singled out or their records linked even if their names remain unknown.
08 · Frequently Asked Questions

Frequently Asked Questions About Identification Without Names

Can someone be identifiable if nobody knows their name?

Yes. A person may be singled out or distinguished from others even before their civil identity is known. Additional information may later connect that distinctive record to a named person.

Can age and occupation identify someone?

Potentially. Neither may be identifying alone, but age and occupation combined with location, sex or gender, institutional role, or another characteristic may substantially narrow the possible individuals.

Can a research quotation identify a participant?

Yes. A quotation may mention a distinctive event or circumstance, or people familiar with the participant may recognise the content or context. Removing the speaker's name does not necessarily prevent recognition.

Can public information make research data identifiable?

Yes. Public directories, institutional websites, professional profiles, news reports, government records, or other datasets may contain overlapping information that enables linkage to a research record.

Does replacing a name with a participant number prevent identification?

Not if the number can be connected to the participant using a code key or another information source. Coding can reduce routine exposure of identity without necessarily making the data anonymous.

Are small studies more vulnerable to identification?

They can be. When the underlying population or sample is small, combinations of otherwise ordinary characteristics may correspond to very few people. Context and reporting detail matter as much as the numerical sample size itself.

09 · The Bottom Line

You Do Not Need a Name to Identify a Person

The Bottom Line

A research participant can be identifiable without their name appearing anywhere in the dataset because identity may emerge through singling out, combinations of attributes, contextual recognition, codes, metadata, or linkage with other information.

When assessing identifiability, look beyond obvious identifiers and ask what the remaining information allows a realistic recipient to distinguish, infer, or connect. Removing names is useful, but it is the beginning of that assessment rather than its conclusion.

10 · Sources and Further Reading

Authoritative Sources on Identification Without Names

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes