Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Can a Small Research Sample Make Participants Identifiable Even When Names Are Removed?

A small sample can increase identification risk because fewer people may match each combination of characteristics. There is no universal sample-size threshold, however; identifiability depends on uniqueness, context, reporting detail, and what others already know.

289
Small Samples and Participant Identification Guide 289 of 398
01 · The Question

If There Are Only a Few Participants, Can Readers Work Out Who They Are?

A qualitative study includes eight department chairs. Names are removed. Institutions are given pseudonyms. Participants are labelled P1 through P8.

One quotation is attributed to "the only female chair with more than 25 years at the institution."

Technically, no name appears. Practically, colleagues may need approximately three seconds.

Small samples can make ordinary descriptive details unusually powerful because fewer people satisfy each combination. But sample size alone does not determine identifiability, and there is no universal number below which confidentiality suddenly disappears.

02 · The Short Answer

Yes, but Small Sample Size Is a Risk Factor Rather Than an Automatic Identification Rule

In Brief

A small research sample can make participants easier to identify because demographic characteristics, roles, quotations, events, or other details may correspond to only one or a few people even after names are removed.

There is no universal safe or unsafe sample size. Identification risk depends on the underlying population, uniqueness of participants, richness of the data, combinations of variables, audience knowledge, reporting practices, and other information available for linkage.

03 · What You Need to Know

Small Samples Increase Risk When They Make People Easier to Single Out

The Problem Is Uniqueness, Not the Number Five, Ten, or Twenty

Researchers sometimes look for a numerical rule: below what sample size do data become identifiable?

There is no universal threshold.

A sample of five participants drawn anonymously from millions of people may reveal very little about who they are. A sample of 100 people consisting of almost every member of one highly visible professional group may still contain distinctive individuals.

The relevant question is whether information allows a person to be singled out or linked to other information. The ICO identifies singling out and linkability as key indicators of identifiability and emphasizes that the risk depends on context and the richness of the data.

Small Samples Shrink the Set of Possible People

Suppose a national survey reports that one participant is a 52-year-old male professor. Thousands of people may match.

Now suppose the study involves six professors from one small department. If only one is male and approximately 52, the same information becomes far more revealing.

Nothing about "52" or "male" changed intrinsically. The comparison population changed.

This illustrates why identifiability is contextual. A year of birth may distinguish someone within a family while doing little to distinguish them within a large school cohort, an example used in current ICO guidance.

A Small Sample Can Turn Indirect Identifiers Into Near-Direct Identifiers

Age, occupation, gender, academic rank, location, years of service, educational background, and similar variables are often indirect identifiers. In large populations, many people may share each value.

Within a small sample, combinations become sparse.

For example:

  • female;
  • full professor;
  • engineering;
  • over 60;
  • more than 30 years at the institution.

If only one person fits that description, the combination effectively singles her out.

This is why researchers should examine how demographic details interact rather than evaluating each characteristic independently.

The Underlying Population Can Matter More Than the Sample

A study may recruit only ten participants but draw them from a population of thousands with no public sampling frame. Another may recruit ten of the twelve people holding a particular position in one organisation.

The second situation can create substantially more recognition risk because readers may already know the small universe from which participants were drawn.

Researchers should therefore distinguish the sample size from the size and visibility of the source population.

Small sample The research includes relatively few participants.
Small identifiable population The possible people from whom participants could have been drawn are few, distinctive, or already known to likely readers.

The two often coincide, but they are not the same problem.

Sampling Information Can Reveal More Than Researchers Expect

A paper may carefully remove participant names while describing recruitment in enough detail to reconstruct the participant pool.

Statements such as "all four female deans at University X participated" effectively identify membership in the sample even if subsequent quotations use participant codes.

Likewise, saying that "three of the institution's four vice presidents participated" sharply narrows the possible speakers whenever a quotation is attributed to a vice president.

Confidentiality review should therefore include the Methods section, sample table, appendices, supplementary files, and other contextual information, not only quotations and raw data.

Quotations Can Be More Identifying Than Demographic Tables

Qualitative quotations contain vocabulary, events, relationships, opinions, and biographical information that may be recognisable to people familiar with participants.

A quotation such as "when I became dean immediately after the 2024 merger" may identify someone more readily than an age or gender variable.

HHS advisory materials have recognised that qualitative recordings may be identifiable depending on circumstances such as a limited sample and unique voice characteristics.

OHRP's exploratory work on qualitative research also notes risks from sensitive or identifiable participant and third-party information and discusses measures such as redaction, site pseudonyms, and attention to dissemination.

Participants May Recognise Themselves and One Another

A small-sample report may permit participants to identify other participants even when outside readers cannot.

This can be particularly important in research involving illegal or socially sensitive behaviour. SACHRP has noted that in small studies, returning general results can sometimes create a possibility that participants identify themselves or other participants.

The likely audience therefore matters. Researchers should consider recognition by participants, colleagues, employers, community members, and other insiders, not merely an anonymous public reader.

Reporting Percentages Can Accidentally Reveal Individuals

Small samples create a less obvious disclosure problem: percentages can look aggregated while still representing very few people.

If a study contains eight participants, 12.5% represents one participant. Reporting that "12.5% of participants disclosed condition X" may therefore communicate that exactly one person did so.

If the paper also provides demographic breakdowns or quotations, readers may be able to determine which participant that was.

Aggregation is useful only when the resulting groups are sufficiently non-distinctive for the intended release context.

Cross-Tabulation Can Create Tiny Cells

A table may look safe because no names appear. But intersecting categories can create cells containing one or two people.

For example:

Rank Male Female Total
Professor 4 1 5
Associate Professor 2 3 5
Assistant Professor 0 4 4

If readers know the institution has only one female professor in the relevant population, a result reported specifically for that cell may effectively become participant-level information.

There is no universal minimum cell size that guarantees anonymity across all research contexts. Statistical agencies, repositories, institutions, and regulatory frameworks may impose specific suppression rules, but researchers should use the rule governing their dataset rather than inventing one.

Removing Demographics Can Reduce Risk but Also Damage the Research

The solution is not necessarily to delete every participant characteristic.

Demographics may be necessary to interpret transferability, describe inequities, examine subgroup differences, or understand the phenomenon being studied. Qualitative research may depend on role and context for interpretation.

Researchers need to balance disclosure risk with scientific meaning. Options may include broader categories, suppression of particular combinations, less precise site descriptions, controlled access, modified attribution of quotations, or withholding details only where they create disproportionate identification risk.

The objective is not to make every participant indistinguishable from every human being on Earth. It is to manage realistic identification risk without quietly deleting the variables that made the study worth conducting.

Combining Sections of a Paper Can Reconstruct Identity

Disclosure risk often emerges across the article rather than within one table.

The Methods section identifies the institutions. Table 1 provides age, gender, and rank. The Results section attributes quotations by participant code. An appendix lists years of experience. Each element may appear acceptable alone.

Together, they may provide enough information to reconstruct identities.

This is essentially the same linkage problem that occurs when multiple datasets are combined. In publication, the "datasets" may simply be different tables, quotations, and contextual descriptions within the same paper.

Watch Out

Do not assess disclosure risk one table or quotation at a time. A reader sees the entire paper and may also know the study setting. Details that appear harmless separately can become identifying when combined.

Small Samples Do Not Automatically Make Confidentiality Impossible

Small-sample research is essential in case studies, rare populations, specialist professions, qualitative inquiry, pilot studies, organisational research, and many other designs.

The appropriate response is not to declare such research inherently non-confidential. It is to recognise that confidentiality may require more contextual judgment than simply removing names.

Factor Lower Identification Concern Higher Identification Concern
Source population Large and not readily known Small, bounded, or publicly known
Participant characteristics Common, broad categories Rare or unique combinations
Geography or institution Broadly described Precisely identified small setting
Quotations Non-distinctive content Unique events, language, or biographical clues
Reporting Groups remain sufficiently non-distinctive Single-person or tiny cells
Audience knowledge Little contextual knowledge Colleagues, community members, or other insiders
04 · A Practical Example

How a Sample of Eight Can Reveal More Than the Researcher Intended

Hypothetical Example

Interviews With Academic Deans

A researcher interviews eight deans from one university about leadership challenges. Names are removed and each participant receives a code.

Methods section The paper states that all eight deans at the university participated.
Participant table The table shows gender, discipline, years in office, and age group.
Results section A quotation from P6 discusses becoming dean during a widely publicised institutional merger.
Reader knowledge University employees know which dean was appointed during the merger and can match the person's discipline and demographics to the participant table.
Disclosure review The researcher revises unnecessary identifying combinations and quotation attribution while retaining contextual information needed to interpret the findings.

No single element necessarily caused the problem. Identification emerged from the cumulative information available to a reader who already knew the setting.

05 · What Researchers Often Get Wrong

Common Misunderstandings About Small Samples and Confidentiality

Misconception

Any Sample Below Ten Is Identifiable

There is no universal sample-size threshold. Identification depends on the population, variables, uniqueness, context, audience knowledge, and release environment rather than sample size alone.

Misconception

A Participant Code Protects People in a Small Sample

Codes remove direct labels but cannot prevent recognition from unique roles, demographic combinations, quotations, or events. P1 can be perfectly obvious without ever becoming Professor Reyes on the page.

Misconception

Aggregate Results Cannot Identify Individuals

Small cells and percentages representing one or two participants can disclose participant-level information, particularly when combined with demographic or contextual information.

Misconception

Removing All Demographics Is the Safest Solution

It may reduce identification risk but can also undermine interpretation and valid subgroup analysis. Researchers should remove, generalise, or restrict information according to actual disclosure risk and scientific necessity.

Misconception

If Journal Readers Cannot Identify Participants, Confidentiality Is Protected

Participants, colleagues, employers, community members, or others familiar with the setting may recognise people from information that means little to outsiders.

06 · What This Means for You

Assess Uniqueness, Not Merely Sample Size

When working with a small sample, examine how much each piece of information narrows the possible participants and how the details interact across the study.

A simple decision framework

If the source population itself is small or publicly known
Assume likely readers may already know the possible participants and assess disclosure from that perspective.
If a demographic combination describes only one or very few participants
Consider whether all variables and their current precision are necessary in the released output.
If quotations contain distinctive events or language
Review contextual recognition risk rather than relying on removal of the speaker's name.
If tables contain cells representing one or very few people
Apply any required disclosure-control rules and consider suppression, grouping, or another suitable reporting approach.
If useful detail cannot be released publicly without unacceptable identification risk
Consider whether controlled access or a less detailed public output better preserves both research value and confidentiality.

A small sample is not a confidentiality failure waiting to happen. It simply leaves less statistical camouflage, so every descriptive detail has to work a little harder for its place on the page.

07 · A Quick Checklist

Check Small Samples for Contextual Identification

Before sharing or publishing findings from a small sample, check:
Consider the size and visibility of the underlying population, not only the numerical sample.
Identify participants with rare roles, characteristics, experiences, or combinations of demographics.
Review participant tables and cross-tabulations for cells containing one or very few people.
Translate percentages back into participant counts to see whether apparently aggregated findings disclose individuals.
Review quotations for distinctive events, language, roles, and biographical clues.
Assess identification using all information across the paper, report, appendices, and supplementary materials together.
Consider whether participants, colleagues, employers, or community insiders could recognise people even when outsiders cannot.
Apply any specific cell-size, suppression, disclosure-control, or repository requirements governing the dataset.
08 · Frequently Asked Questions

Frequently Asked Questions About Small Samples and Identification

What sample size is too small for anonymity?

There is no universal number. Identifiability depends on how distinctive participants are, the size of the source population, what information is released, who receives it, and what additional information is available.

Is a qualitative sample of ten automatically identifiable?

No. A sample of ten can present low or high identification risk depending on the setting and information. Rich qualitative narratives and bounded populations may increase contextual recognition risk, but the number ten itself does not determine identifiability.

Can percentages identify participants in a small sample?

Yes. In small samples, percentages may correspond to one or very few people. Combined with demographic or contextual information, apparently aggregate statistics can reveal participant-level information.

Should I remove the participant characteristics table from a small qualitative study?

Not automatically. Participant characteristics may be important for interpretation. Review whether the particular variables, precision, combinations, and attribution create disproportionate identification risk and modify only what is necessary.

Can participants identify each other even if the public cannot?

Yes. Participants often possess contextual information unavailable to outside readers. Shared workplaces, communities, events, or experiences can make apparently anonymous descriptions recognisable.

Does using pseudonyms solve small-sample identification?

No. Pseudonyms remove names but not unique roles, demographic combinations, distinctive quotations, events, or contextual knowledge that can reveal identity.

09 · The Bottom Line

Small Samples Make Context More Powerful, Not Confidentiality Impossible

The Bottom Line

A small sample can make participants identifiable even without names because fewer people may match each combination of characteristics, roles, events, quotations, or reported results.

There is no universal sample-size cutoff. Assess the underlying population, uniqueness, data richness, audience knowledge, small cells, and cumulative information released across the study. In small samples, context often does the identifying work that a name would otherwise do.

10 · Sources and Further Reading

Authoritative Sources on Small-Sample Identifiability

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes