Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Counts as Identifiable Research Information?

Identifiable research information includes more than names and contact details. Participants may be identifiable directly, through codes, from combinations of characteristics, or by linking research records with other available information.

275
Identifiable Research Information Guide 275 of 398
01 · The Question

When Does Research Information Become Identifiable?

A spreadsheet contains no participant names. Does that make it non-identifiable?

Not necessarily. A participant might be identified by an email address, photograph, voice, student number, precise location, unusual job title, combination of demographic characteristics, or a code that connects the record to another file. Sometimes no single field identifies anyone, yet several fields together point to one person.

Researchers therefore need a broader test than "Did I collect names?" The real issue is whether a person's identity is known, associated with the information, or can reasonably be determined under the standard that applies to the research.

02 · The Short Answer

Information Is Identifiable When It Reveals or Can Be Connected to a Person

In Brief

Research information is identifiable when it directly identifies a person or when identity can be determined by connecting the information with codes, other variables, contextual knowledge, or additional reasonably available information under the applicable legal or ethical standard.

Names are only the obvious case. Contact details, identification numbers, images, voices, precise dates and locations, distinctive characteristics, combinations of demographic variables, free-text responses, metadata, and linked datasets may also make research information identifiable.

03 · What You Need to Know

Identifiability Extends Far Beyond a Participant's Name

Some Information Identifies People Directly

Direct identifiers point to a particular person with little or no additional inference. Common examples can include a person's full name, personal email address, telephone number, government-issued identification number, student or employee number, account identifier, or other unique personal identifier.

Some forms of biometric or multimedia information may also identify a person directly or very strongly. HIPAA's Safe Harbor provisions, for example, include biometric identifiers such as finger and voice prints, full-face photographs and comparable images, Internet Protocol addresses, URLs, medical record numbers, and several other categories among the identifiers addressed by its specific de-identification method.

That list is useful for understanding how varied identifying information can be, but it should not be treated as a universal checklist for all research. HIPAA applies to particular health information and entities in the United States. Other laws and research frameworks define identifiability differently.

Information Can Identify Someone Indirectly

Indirect identifiers do not necessarily reveal identity by themselves. They narrow the field of possible people or become identifying when combined with other information.

Examples may include age, sex or gender, occupation, institutional role, education, geographic location, dates, nationality, household composition, uncommon characteristics, or combinations of demographic details.

Whether a variable identifies someone depends heavily on context. "Age 40" may describe thousands of people. "Age 40, the only female neurosurgeon at a particular small hospital" may narrow the possibilities dramatically.

This is why the distinction between direct and indirect identifiers is more useful than simply maintaining a list of forbidden columns.

A Combination Can Identify Someone Even When No Variable Does Alone

Identifiability can emerge from combinations. Imagine a dataset containing only age, municipality, occupation, and highest educational qualification. Each field may look ordinary. Together, a rare combination might correspond to only one person in the relevant population.

This becomes particularly important in small or specialised research populations. Detailed demographic reporting that is harmless in a national survey may become revealing in a study of a small department, rare profession, tiny community, or distinctive patient group.

Researchers should therefore evaluate the joint information content of the dataset rather than reviewing variables one at a time. Demographic details can become identifying through their combination and context.

A Participant Can Be Identifiable Even If You Do Not Know Their Name

Identification and naming are not identical. A dataset might allow a researcher to single out one particular individual repeatedly without revealing that person's civil identity. Another party may then be able to connect that record to a real-world identity using additional information.

Similarly, a distinctive description may reveal exactly which colleague, patient, student, public official, or community member is being discussed to people familiar with the setting, even if an outside reader would not know the person's name.

Researchers should therefore ask whether information can identify someone without explicitly naming them.

Coded Data May Still Be Identifiable

Replacing "Ana Reyes" with "P103" can reduce exposure of identity, but the status of the information depends partly on whether the code can be resolved.

The US Office for Human Research Protections defines coded information for its guidance as information in which identifying information has been replaced with a code and a key exists that enables linkage. OHRP generally considers private information identifiable for Common Rule purposes when investigators can link it to specific individuals directly or indirectly through coding systems.

However, OHRP also explains that in certain secondary-research circumstances coded information may not be considered individually identifiable to the investigator when the investigator cannot readily ascertain identities, such as where an agreement or repository policy prevents release of the key.

This is a useful reminder that identifiability may depend on who holds the information and what additional information that person can access.

Free-Text Responses Can Contain Identifiers Researchers Never Asked For

Structured questionnaires make it relatively easy to see what fields are being collected. Free text is less cooperative.

A participant asked to describe a workplace experience might write: "When I became dean in 2024, I was the first person from our department appointed to the position." Another might name colleagues, organisations, locations, illnesses, court cases, or distinctive events.

Researchers therefore cannot assume that a survey is non-identifiable merely because the questionnaire contains no name field. Open-ended responses can introduce identifying information spontaneously.

The same issue arises in interview transcripts, diaries, field notes, case histories, discussion-board posts, and other narrative materials.

Images, Audio, Video, and Biospecimens Need Their Own Assessment

Research information is not limited to spreadsheet variables.

A photograph may contain a participant's face, home, workplace, vehicle plate, identification badge, or location clues. Video can capture faces, voices, surroundings, and interactions. Audio may include a recognisable voice or spoken identifying details.

Biospecimens also raise distinct identifiability questions. Under the current US Common Rule, an identifiable biospecimen is one for which the identity of the subject is or may readily be ascertained by the investigator or associated with the biospecimen. OHRP applies a similar formulation to identifiable private information.

The applicable regulatory treatment may differ across jurisdictions and research contexts, so researchers should not assume that the rules for a spreadsheet automatically settle the status of images, recordings, or biological material.

Metadata Can Identify Participants Without Appearing in the Main Dataset

Identifying information can also sit around the data rather than inside the variables researchers intend to analyse.

Online systems may generate IP addresses, timestamps, device information, usernames, account identifiers, file properties, geolocation data, or unique survey links. Digital photographs may contain metadata. Document files may preserve author names or revision histories.

This means a researcher assessing identifiability should inspect the entire data flow, not just the exported analysis spreadsheet.

Linkage Changes What Counts as Identifying

A variable that seems harmless in isolation can become identifying when matched against another source. For example, an exact date, postcode, occupation, and demographic characteristic might correspond to publicly available professional information.

HIPAA's Expert Determination approach explicitly considers whether information could be used alone or in combination with other reasonably available information by an anticipated recipient to identify an individual.

This illustrates the broader principle behind re-identification through combined datasets: identifiability depends not only on what your file contains, but also on what can realistically be connected to it.

Legal Definitions Are Similar in Purpose but Not Identical

Researchers often work across several frameworks, and their definitions should not be collapsed into one universal test.

Framework Relevant Idea Important Qualification
US Common Rule / OHRP Identifiable private information is private information for which the subject's identity is or may readily be ascertained by the investigator or is associated with the information. The investigator's ability to ascertain identity matters, including linkage through coding systems.
HIPAA Privacy Rule Individually identifiable health information includes information that identifies an individual or for which there is a reasonable basis to believe it can identify the individual. HIPAA applies to particular health information held or transmitted by covered entities and business associates, not all research data.
GDPR-style data protection Personal data concern an identified or identifiable natural person, with identifiability assessed using relevant means of identification. Pseudonymized information can remain personal data when additional information enables attribution.

The practical lesson is not to choose whichever definition produces the easiest answer. Determine which law, ethics framework, institutional policy, repository requirement, or contractual obligation applies to the research and use its terminology consistently.

04 · A Practical Example

How an Apparently Nameless Dataset Can Still Identify a Participant

Hypothetical Example

A Study of University Department Chairs

A researcher interviews department chairs at several universities and removes every participant's name before creating the analysis dataset.

Direct identifiers removed Names, email addresses, and telephone numbers no longer appear in the transcript.
Indirect information remains One transcript describes the participant as a 38-year-old woman who chairs a particular specialist discipline at the only university offering that program in her province.
Free text adds context The participant mentions receiving a publicly announced national award the previous year.
External information completes the link The university website lists the department chair, while a news release identifies the award recipient.
Result Removing the name reduced direct identification, but the remaining combination may still make the participant readily identifiable to someone with relevant contextual information.

The important question was never merely whether the transcript contained a name. It was whether the information remaining in and around the transcript could reasonably be associated with a particular person.

05 · What Researchers Often Get Wrong

Common Mistakes When Deciding Whether Research Data Are Identifiable

Misconception

Only Names Count as Identifiers

Names are merely the most obvious example. Contact information, identification numbers, images, voices, account information, precise locations, distinctive characteristics, metadata, and combinations of ordinary variables may also identify participants.

Misconception

Every Demographic Variable Is Automatically Identifying

Not necessarily. Identifiability depends on granularity, combinations, population size, context, and available external information. A broad age category may carry little identification risk in one dataset and contribute strongly to identification in another.

Misconception

A Participant Code Makes Information Non-Identifiable

If the investigator can use a key to connect the code to the participant, the information may remain identifiable under the applicable framework. Coding can still be a valuable safeguard because it reduces routine exposure of direct identifiers.

Misconception

If I Cannot Recognise the Participant, the Data Are Non-Identifiable

Your own knowledge is not always the entire test. Another researcher, data recipient, insider, or member of the community may possess additional information that makes identification possible. The relevant standard determines whose capabilities and information must be considered.

Misconception

Only the Variables I Plan to Analyse Matter

Identifying information may exist in consent records, contact files, code keys, metadata, filenames, audio recordings, photographs, free text, system logs, or other accompanying materials. Assess the complete research information environment.

06 · What This Means for You

Map Identifiability Before Deciding How to Protect the Data

Identifiability should be assessed during study design, not discovered after data collection. Start by asking what information is genuinely necessary for the research and what each additional detail allows someone to infer.

A simple decision framework

If a variable directly identifies a participant
Determine whether it is necessary and, if retained, restrict access and protect it appropriately.
If several variables become identifying only when combined
Consider reducing granularity, suppressing unnecessary detail, or applying other appropriate disclosure controls while preserving necessary research utility.
If identifying information is needed only for recruitment or follow-up
Consider separating participant identifiers from research data and limiting access to the linkage information.
If open-ended or multimedia data are collected
Plan specifically for identifying content within narratives, images, audio, video, and metadata rather than relying only on structured-field controls.
If you intend to share or publish data
Reassess identifiability from the recipient's perspective and in the intended release environment.

Collecting fewer unnecessary identifiers can simplify almost every later stage of data governance. If information serves no defensible research or operational purpose, the safest identifier may be the one that was never collected in the first place.

07 · A Quick Checklist

Check More Than the Name Column

When assessing whether research information is identifiable, check:
List direct identifiers such as names, contact details, identification numbers, account identifiers, and other unique identifiers.
Review demographic and contextual variables individually and in combination.
Determine whether participant codes can be resolved using a key or other available information.
Inspect free-text responses, transcripts, field notes, and quotations for names and distinctive contextual details.
Assess photographs, video, audio, biometrics, and biospecimens using standards appropriate to those data types.
Check metadata, filenames, timestamps, IP addresses, unique links, account information, and system-generated records.
Consider whether external datasets or publicly available information could be linked to the research records.
Assess identifiability from the perspective of relevant researchers, recipients, insiders, or other parties required by the applicable framework.
Verify the definition of identifiable information used by your ethics body, institution, applicable law, and data-sharing agreement.
08 · Frequently Asked Questions

Frequently Asked Questions About Identifiable Research Information

Is a participant's name always identifiable information?

A name ordinarily functions as a direct identifier when it connects research information to a particular person, although common names may not uniquely identify someone without context. Researchers should still treat names as potentially identifying and apply the relevant requirements.

Can age identify a research participant?

It can contribute to identification, particularly when exact age is combined with location, occupation, sex or gender, rare characteristics, or a small study population. Its identifying power depends on context and granularity.

Is a participant ID number identifiable information?

It may be. If the number can be connected to an individual using a code key, administrative database, or other information, the record may remain identifiable to parties able to make that connection.

Can a quotation identify a participant?

Yes. A quotation may contain names or distinctive details, or the wording itself may be recognisable to people familiar with the speaker or event. Qualitative researchers should assess quotations before dissemination rather than assuming removal of names is sufficient.

Is an IP address identifying information?

It can be identifying or personal information depending on the applicable framework and circumstances. HIPAA Safe Harbor, for example, includes IP addresses among the identifiers specified for removal under that particular de-identification method.

Can someone be identifiable even if I do not know their name?

Yes. Information may single out or distinguish a particular person, and another party may be able to connect that information to a real-world identity using contextual knowledge or additional data.

Does removing direct identifiers make a dataset anonymous?

Not automatically. Indirect identifiers, combinations of variables, code keys, free text, metadata, and external linkage may still permit identification. Anonymity requires a broader assessment of realistic identifiability.

09 · The Bottom Line

Identifiability Is About Connections, Not Just Names

The Bottom Line

Research information is identifiable when it directly reveals a person's identity or when identity can be determined through codes, combinations of characteristics, contextual knowledge, metadata, or linkage with other information under the applicable standard.

Review the entire information environment rather than searching only for obvious identifiers. What looks harmless in isolation can become identifying in combination, particularly in small populations or when detailed external information is available.

10 · Sources and Further Reading

Authoritative Sources on Identifiable Research Information

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes