03 · What You Need to Know
Identifiability Extends Far Beyond a Participant's Name
Some Information Identifies People Directly
Direct identifiers point to a particular person with little or no additional inference. Common examples can include a person's full name, personal email address, telephone number, government-issued identification number, student or employee number, account identifier, or other unique personal identifier.
Some forms of biometric or multimedia information may also identify a person directly or very strongly. HIPAA's Safe Harbor provisions, for example, include biometric identifiers such as finger and voice prints, full-face photographs and comparable images, Internet Protocol addresses, URLs, medical record numbers, and several other categories among the identifiers addressed by its specific de-identification method.
That list is useful for understanding how varied identifying information can be, but it should not be treated as a universal checklist for all research. HIPAA applies to particular health information and entities in the United States. Other laws and research frameworks define identifiability differently.
Information Can Identify Someone Indirectly
Indirect identifiers do not necessarily reveal identity by themselves. They narrow the field of possible people or become identifying when combined with other information.
Examples may include age, sex or gender, occupation, institutional role, education, geographic location, dates, nationality, household composition, uncommon characteristics, or combinations of demographic details.
Whether a variable identifies someone depends heavily on context. "Age 40" may describe thousands of people. "Age 40, the only female neurosurgeon at a particular small hospital" may narrow the possibilities dramatically.
This is why the distinction between direct and indirect identifiers is more useful than simply maintaining a list of forbidden columns.
A Combination Can Identify Someone Even When No Variable Does Alone
Identifiability can emerge from combinations. Imagine a dataset containing only age, municipality, occupation, and highest educational qualification. Each field may look ordinary. Together, a rare combination might correspond to only one person in the relevant population.
This becomes particularly important in small or specialised research populations. Detailed demographic reporting that is harmless in a national survey may become revealing in a study of a small department, rare profession, tiny community, or distinctive patient group.
Researchers should therefore evaluate the joint information content of the dataset rather than reviewing variables one at a time. Demographic details can become identifying through their combination and context.
A Participant Can Be Identifiable Even If You Do Not Know Their Name
Identification and naming are not identical. A dataset might allow a researcher to single out one particular individual repeatedly without revealing that person's civil identity. Another party may then be able to connect that record to a real-world identity using additional information.
Similarly, a distinctive description may reveal exactly which colleague, patient, student, public official, or community member is being discussed to people familiar with the setting, even if an outside reader would not know the person's name.
Researchers should therefore ask whether information can identify someone without explicitly naming them.
Coded Data May Still Be Identifiable
Replacing "Ana Reyes" with "P103" can reduce exposure of identity, but the status of the information depends partly on whether the code can be resolved.
The US Office for Human Research Protections defines coded information for its guidance as information in which identifying information has been replaced with a code and a key exists that enables linkage. OHRP generally considers private information identifiable for Common Rule purposes when investigators can link it to specific individuals directly or indirectly through coding systems.
However, OHRP also explains that in certain secondary-research circumstances coded information may not be considered individually identifiable to the investigator when the investigator cannot readily ascertain identities, such as where an agreement or repository policy prevents release of the key.
This is a useful reminder that identifiability may depend on who holds the information and what additional information that person can access.
Free-Text Responses Can Contain Identifiers Researchers Never Asked For
Structured questionnaires make it relatively easy to see what fields are being collected. Free text is less cooperative.
A participant asked to describe a workplace experience might write: "When I became dean in 2024, I was the first person from our department appointed to the position." Another might name colleagues, organisations, locations, illnesses, court cases, or distinctive events.
Researchers therefore cannot assume that a survey is non-identifiable merely because the questionnaire contains no name field. Open-ended responses can introduce identifying information spontaneously.
The same issue arises in interview transcripts, diaries, field notes, case histories, discussion-board posts, and other narrative materials.
Images, Audio, Video, and Biospecimens Need Their Own Assessment
Research information is not limited to spreadsheet variables.
A photograph may contain a participant's face, home, workplace, vehicle plate, identification badge, or location clues. Video can capture faces, voices, surroundings, and interactions. Audio may include a recognisable voice or spoken identifying details.
Biospecimens also raise distinct identifiability questions. Under the current US Common Rule, an identifiable biospecimen is one for which the identity of the subject is or may readily be ascertained by the investigator or associated with the biospecimen. OHRP applies a similar formulation to identifiable private information.
The applicable regulatory treatment may differ across jurisdictions and research contexts, so researchers should not assume that the rules for a spreadsheet automatically settle the status of images, recordings, or biological material.
Metadata Can Identify Participants Without Appearing in the Main Dataset
Identifying information can also sit around the data rather than inside the variables researchers intend to analyse.
Online systems may generate IP addresses, timestamps, device information, usernames, account identifiers, file properties, geolocation data, or unique survey links. Digital photographs may contain metadata. Document files may preserve author names or revision histories.
This means a researcher assessing identifiability should inspect the entire data flow, not just the exported analysis spreadsheet.
Linkage Changes What Counts as Identifying
A variable that seems harmless in isolation can become identifying when matched against another source. For example, an exact date, postcode, occupation, and demographic characteristic might correspond to publicly available professional information.
HIPAA's Expert Determination approach explicitly considers whether information could be used alone or in combination with other reasonably available information by an anticipated recipient to identify an individual.
This illustrates the broader principle behind re-identification through combined datasets: identifiability depends not only on what your file contains, but also on what can realistically be connected to it.
Legal Definitions Are Similar in Purpose but Not Identical
Researchers often work across several frameworks, and their definitions should not be collapsed into one universal test.
| Framework |
Relevant Idea |
Important Qualification |
|
US Common Rule / OHRP
|
Identifiable private information is private information for which the subject's identity is or may readily be ascertained by the investigator or is associated with the information. |
The investigator's ability to ascertain identity matters, including linkage through coding systems. |
|
HIPAA Privacy Rule
|
Individually identifiable health information includes information that identifies an individual or for which there is a reasonable basis to believe it can identify the individual. |
HIPAA applies to particular health information held or transmitted by covered entities and business associates, not all research data. |
|
GDPR-style data protection
|
Personal data concern an identified or identifiable natural person, with identifiability assessed using relevant means of identification. |
Pseudonymized information can remain personal data when additional information enables attribution. |
The practical lesson is not to choose whichever definition produces the easiest answer. Determine which law, ethics framework, institutional policy, repository requirement, or contractual obligation applies to the research and use its terminology consistently.
07 · A Quick Checklist
Check More Than the Name Column
When assessing whether research information is identifiable, check:
List direct identifiers such as names, contact details, identification numbers, account identifiers, and other unique identifiers.
Review demographic and contextual variables individually and in combination.
Determine whether participant codes can be resolved using a key or other available information.
Inspect free-text responses, transcripts, field notes, and quotations for names and distinctive contextual details.
Assess photographs, video, audio, biometrics, and biospecimens using standards appropriate to those data types.
Check metadata, filenames, timestamps, IP addresses, unique links, account information, and system-generated records.
Consider whether external datasets or publicly available information could be linked to the research records.
Assess identifiability from the perspective of relevant researchers, recipients, insiders, or other parties required by the applicable framework.
Verify the definition of identifiable information used by your ethics body, institution, applicable law, and data-sharing agreement.