Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How Much Identifying Information Should Researchers Collect From Participants?

Researchers should not collect identifying information simply because it might be useful later. Each identifier should serve a defensible research or operational purpose, and the appropriate amount depends on what the study actually needs to accomplish.

280
How Much Identifying Information Should You Collect? Guide 280 of 398
01 · The Question

How Do You Know Whether You Are Collecting Too Much Identifying Information?

A registration form asks for a participant's full name, email address, telephone number, exact age, department, job title, and location. Each field seems potentially useful. Taken together, however, they create a detailed identity profile.

Researchers often focus on protecting identifiers after collection. An earlier question may be even more useful: did the study need to collect each identifier in the first place?

The appropriate amount is not always zero. Follow-up studies, longitudinal designs, incentive distribution, clinical procedures, record linkage, withdrawal mechanisms, and other legitimate research functions may require identifying information. The task is to distinguish what the study needs from what is merely convenient to have.

02 · The Short Answer

Collect What the Study Needs, but No More

In Brief

Researchers should collect the minimum identifying information necessary to accomplish clearly defined research and legitimate operational purposes, rather than collecting identifiers speculatively because they might become useful later.

What counts as necessary depends on the study. For each identifier, researchers should be able to explain why it is needed, whether a less identifying alternative would work, how precisely it must be collected, who needs access, and when the identifying information is no longer necessary.

03 · What You Need to Know

Data Minimisation Starts Before the First Participant Enrols

Begin With Purpose, Not With a Standard Demographic Form

The easiest way to collect unnecessary personal information is to begin with last year's questionnaire and leave every field in place.

A better approach starts with purpose. For every item that identifies or helps identify a participant, ask what function it serves in this particular study.

A name might be needed for consent documentation or participant communication. An email address might be necessary for follow-up. A telephone number might be unnecessary if all communication occurs by email. Exact date of birth might be excessive if the analysis needs only broad age categories.

Under the UK GDPR's data-minimisation principle, for example, personal data should be adequate, relevant, and limited to what is necessary for the purposes for which they are processed. ICO guidance translates this into a practical instruction: identify the minimum amount of personal data needed to fulfil the purpose and hold that much, but no more.

That is a useful design question even where UK law does not apply. The applicable legal obligations will vary by jurisdiction, but unnecessary identifiers generally create privacy and governance burdens without improving the science.

"Necessary" Does Not Mean "Might Be Useful Someday"

Research produces uncertainty, and researchers understandably want flexibility. That does not make every potentially useful variable necessary.

The ICO's research guidance makes a useful distinction: necessary processing need not be absolutely indispensable, but it should be a targeted and proportionate way of achieving the research purpose. If the same purpose can reasonably be achieved through less intrusive means, collecting additional personal data requires stronger justification.

For example, collecting participants' personal mobile numbers because researchers might decide to conduct follow-up interviews later is weaker justification than collecting them because follow-up interviews are an approved component of the protocol.

The difference is between a defined purpose and speculative convenience.

Collecting No Identifier at All May Sometimes Be Possible

If a study does not require follow-up, record linkage, verification, incentives tied to identity, or another participant-specific function, researchers should consider whether identifying information is needed at all.

ICO research guidance explicitly advises researchers first to consider whether research can be conducted without personal data and, where possible, using anonymous information.

This does not mean every study should be anonymous. Many cannot be. It means researchers should not assume that identity is a default variable.

The distinction between anonymous and confidential research becomes important here. If identity is genuinely unnecessary, designing the study so the researcher never receives it may eliminate an entire category of confidentiality risk.

Collect the Least Identifying Version That Answers the Research Question

Data minimisation is not only about whether a variable is collected. It also concerns precision.

Suppose age matters analytically. Do you need:

  • date of birth;
  • exact age;
  • five-year age bands; or
  • broad age categories?

The scientifically appropriate answer depends on the analysis. A study examining age-related developmental change may legitimately require greater precision than a study using age only to describe the sample.

The same question applies to geography. An exact home address may be essential for environmental exposure modelling but excessive for a study that needs only region of residence.

Data minimisation therefore does not mean mechanically choosing the least detailed variable. The data must remain adequate for the research purpose. The objective is the least identifying level of detail that still permits the study to do what it legitimately needs to do.

Distinguish Research Variables From Administrative Identifiers

Some identifying information is analytically important. Other information exists only to operate the study.

Research information Information needed to answer the research question, test hypotheses, describe the sample, control confounding, conduct planned analyses, or otherwise achieve the scientific purpose.
Administrative information Information needed for recruitment, scheduling, follow-up, incentives, consent administration, withdrawal requests, or other study operations but not necessarily for analysis.

This distinction matters because administrative identifiers often do not need to travel with the research dataset. If email addresses are needed only to schedule interviews, there may be little reason for analysts to receive them alongside interview responses.

Researchers can therefore ask not only whether an identifier must be collected, but also whether it must be stored with the substantive research data.

Every Additional Identifier Expands the Identification Surface

An identifier does not have to be a name to increase identifiability. Exact age, occupation, geographic location, institutional role, event dates, and demographic details may narrow the possible participants.

Collecting several such variables can create distinctive combinations even when none identifies anyone alone. This is why researchers should consider direct and indirect identifiers together.

A useful question is therefore not merely, "Is this field sensitive?" It is also, "What does this field add to the possibility of identifying someone when combined with everything else I collect?"

Small Samples Need Particular Restraint

Detailed demographics may be scientifically useful, but their identifying power changes with population size.

Recording exact age, specialised occupation, department, seniority, gender, and institution may be unremarkable in a national survey involving thousands of participants. The same variables in a study of 12 employees from one organisation may describe individuals almost by name.

Researchers planning small or distinctive samples should therefore examine whether every demographic detail is necessary and whether broader categories would still answer the research question.

This does not justify altering scientifically necessary variables simply to make a dataset look safer. It means identification risk should form part of the study-design decision rather than being discovered at publication.

Collecting an Identifier Creates Obligations Beyond Secure Storage

Once identifying information enters the research system, researchers may need to consider lawful processing, transparency, access control, retention, disclosure, participant expectations, security, and eventual deletion or archival treatment under the rules applicable to their study.

Data minimisation can therefore simplify governance. Information that was never collected cannot be leaked from the research database, accidentally included in an export, unnecessarily shared with collaborators, or retained long after its purpose has ended.

That does not mean "collect nothing." Inadequate data can also compromise research validity. The UK GDPR formulation is deliberately balanced: personal data should be adequate and relevant as well as limited to what is necessary.

More Variables Can Also Create More Re-Identification Opportunities

A detailed dataset may become easier to link with information available elsewhere. Exact dates, locations, job titles, and demographic characteristics can provide matching variables for re-identification through other datasets.

Researchers planning eventual data sharing should therefore consider the downstream consequences of collection choices. It is easier to avoid collecting an unnecessary highly identifying variable than to preserve its full analytical detail later while somehow making its identifying power disappear.

There Is No Universal Maximum Number of Identifiers

Data minimisation is not a numerical quota. Collecting three unnecessary identifiers is not acceptable merely because another study collects ten, and a study requiring several identifiers is not automatically excessive.

The appropriate amount depends on purpose, proportionality, participant population, study procedures, analytical requirements, applicable law, ethics requirements, and available alternatives.

Information Possible Legitimate Purpose Question to Ask
Full name Consent administration, participant-specific follow-up Does the analysis need the name, or can it remain separate?
Email address Scheduling, follow-up, incentive delivery Is email the chosen contact method, and how long must it be retained?
Telephone number Participant contact Is a second contact channel genuinely necessary?
Date of birth Precise age calculation or record linkage Would exact age or an age band accomplish the purpose?
Home address Geospatial exposure analysis or necessary correspondence Is exact location required, or would a broader geographic unit suffice?
Employer or department Sampling, stratification, organisational analysis Does this level of organisational detail materially serve the research question?
Detailed demographics Planned subgroup analysis or confounding control Are the categories analytically justified and proportionate to identification risk?
04 · A Practical Example

From a Convenient Form to a Minimised Data Collection Plan

Hypothetical Example

An Online Survey With Follow-Up Interviews

A researcher plans a survey of university instructors about AI use in teaching. Participants may optionally volunteer for a follow-up interview.

Initial form The draft asks every respondent for full name, university email, mobile number, exact age, institution, college, department, academic rank, years of service, and several research variables.
Return to the research purpose The analysis requires age group, academic rank, discipline, years of teaching experience, and survey responses. Names and contact details are not analytical variables.
Challenge the administrative fields Email addresses are needed only for respondents volunteering for interviews. Mobile numbers are unnecessary because interviews will be scheduled by email.
Reduce precision Exact age is unnecessary because the planned analysis uses age categories. Department name is also unnecessary because a broader discipline category answers the research question.
Redesign collection The main survey collects only the research variables required for analysis. Volunteers provide contact information through a separate process designed around the actual follow-up requirement.

The revised design does not minimise data by deleting useful research variables indiscriminately. It distinguishes analytical necessity from administrative convenience and removes information that serves neither.

05 · What Researchers Often Get Wrong

Common Mistakes When Deciding What Identifying Information to Collect

Misconception

Collect Everything Now Because You Can Always Delete It Later

Later deletion does not undo the privacy and security exposure created while unnecessary information was collected, transmitted, stored, and accessible. Decide whether the information is needed before collecting it whenever possible.

Misconception

Data Minimisation Means Collecting as Little Data as Possible

Not quite. Research data must still be adequate for the intended purpose. Minimisation means avoiding personal data beyond what is necessary, not weakening a valid study by removing variables essential to the research question or analysis.

Misconception

Consent Makes Any Amount of Identifying Information Acceptable

Participant consent to take part does not by itself answer whether every collected identifier is necessary, proportionate, or lawful. Under UK GDPR guidance, ethical consent to participate is also distinct from the lawful basis used to process personal data.

Misconception

Demographic Information Does Not Count Because It Is Not a Name

Detailed demographics can function as indirect identifiers, particularly in small populations. Researchers should justify both the variable and its level of precision.

Misconception

If an Identifier Is Necessary Once, It Must Be Retained for the Whole Project

Different identifiers may be necessary at different stages. Contact information needed during recruitment may not need to remain with analysis data after recruitment ends. Retention requirements depend on the purpose and applicable rules, so researchers should plan the identifier lifecycle rather than assuming permanent necessity.

06 · What This Means for You

Make Every Identifier Earn Its Place in the Dataset

Before finalising a questionnaire, interview form, case-report form, registration system, or data-extraction template, review every field that identifies or helps identify participants.

A simple decision framework

If the study can achieve its purpose without identifying participants
Consider an anonymous design rather than collecting identity by default.
If an identifier is required only for recruitment, scheduling, incentives, or follow-up
Consider collecting and storing it separately from substantive research responses.
If a detailed variable is analytically necessary
Retain the precision the research genuinely requires and document why it is necessary.
If a broader category answers the same research question
Consider collecting the broader information rather than unnecessarily precise personal data.
If you cannot explain why an identifying field is being collected
Do not treat "we might need it later" as sufficient justification without a defined purpose.

A useful protocol table can document each identifying variable, its purpose, level of precision, who needs access, where it is stored, and when that need ends. That small exercise tends to expose unnecessary fields remarkably quickly. Forms, like literature reviews, have a habit of accumulating material nobody remembers adding.

07 · A Quick Checklist

Audit Identifying Information Before Data Collection Begins

For every identifying or potentially identifying variable, check:
State the specific research or operational purpose requiring the information.
Ask whether the same purpose can reasonably be achieved without collecting personal or identifying information.
Use the least precise form of the information that still adequately supports the legitimate research purpose.
Distinguish analytical variables from contact, recruitment, incentive, and other administrative information.
Consider whether administrative identifiers can be collected or stored separately from research responses.
Review combinations of demographic variables for indirect identification risk, particularly in small samples.
Define who needs access to each category of identifying information.
Plan when each identifier will no longer be needed and apply the retention requirements governing the study.
Verify applicable ethics requirements, institutional policies, and data-protection law before data collection begins.
08 · Frequently Asked Questions

Frequently Asked Questions About Collecting Participant Identifiers

Should researchers avoid collecting names whenever possible?

If names serve no legitimate research or operational purpose, avoiding them can reduce identification risk. Some studies legitimately require names for consent administration, follow-up, linkage, participant-specific procedures, or other functions, so the decision depends on the study.

Can I collect contact information in case I want to follow participants up later?

A defined follow-up plan provides stronger justification than collecting contact information speculatively. Researchers should identify the purpose in advance, address it appropriately in the protocol and participant information, and collect only the contact information actually needed.

Should I collect exact age or age groups?

Use the level of precision required by the research. Exact age may be necessary for some analyses, while broad categories may be adequate for others. The decision should reflect analytical requirements as well as identification risk.

Is it safer to collect identifiers and remove them later?

Removing identifiers later can reduce subsequent risk, but the information remains identifiable while those identifiers exist. If an identifier is unnecessary from the outset, not collecting it avoids that period of exposure entirely.

Does data minimisation prevent researchers from collecting rich datasets?

No. It requires a relationship between the information collected and the research purpose. Rich or detailed information may be justified when scientifically necessary; unnecessary personal detail is the concern.

Does data minimisation require deleting research data quickly?

Not necessarily. Collection minimisation and retention are related but distinct questions. Some research frameworks permit or require long-term retention subject to safeguards. Researchers should apply the retention rules and scientific requirements governing their particular study rather than assuming that minimisation means immediate deletion.

09 · The Bottom Line

Collect Identifiers Because You Need Them, Not Because the Form Has Space

The Bottom Line

Researchers should collect only the identifying information necessary for defined research and legitimate operational purposes, at no greater precision than those purposes require.

For each identifier, ask why it is needed, whether a less identifying alternative works, whether it must accompany the research responses, who needs access, and when the need ends. Good data protection begins before information enters the dataset.

10 · Sources and Further Reading

Authoritative Sources on Data Minimisation in Research

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes