01 · The Question
How Much Participant Information Do You Really Need?
Age, sex, gender, email address, occupation, income, exact location, student number, medical history, ethnicity, religion: once researchers start designing a questionnaire or database, the list of potentially useful variables can grow surprisingly quickly.
The temptation is understandable. If you are already collecting data, why not ask a few additional questions in case they become useful later?
That is precisely where researchers should become more selective. Personal data should be connected to a defined research, administrative, safety, or other legitimate purpose. A variable being interesting, conventional, or easy to collect does not by itself make it necessary.
03 · What You Need to Know
Start With Purpose, Then Decide What Data Follow From It
There is no universal list of personal data every study should collect
A longitudinal clinical study, an anonymous classroom survey, an interview study, and a population database analysis have very different information requirements. Asking which personal data researchers are allowed to collect without first asking why the data are needed reverses the decision process.
A better sequence begins with the research question and study design. Determine what needs to be measured, what information is necessary to recruit or communicate with participants, what information is required for safety or follow-up, and what records may be necessary for legitimate administrative or regulatory purposes. Only then should you decide which personal-data fields are needed.
This also means distinguishing the information needed to conduct the study from the information needed to analyze it. A participant's email address might be necessary for scheduling an interview but completely unnecessary in the analytical dataset. Contact details can often be stored separately rather than accompanying the research responses throughout the project.
Personal data are not limited to names and contact details
What legally constitutes personal data depends on the applicable jurisdiction, but researchers should not assume that removing names automatically removes all personal information. Information can identify a person directly or, in some circumstances, indirectly when combined with other information.
A research dataset might therefore contain personal data even when there is no column labeled "Name." Exact dates, detailed geographic information, institutional identifiers, online identifiers, unusual occupations, combinations of demographic characteristics, images, audio recordings, or sufficiently distinctive free-text responses can potentially contribute to identification.
Whether a particular dataset legally constitutes personal data requires assessment under the applicable framework. The practical lesson is simpler: think about identifiability across the dataset, not merely whether obvious identifiers have been removed.
Give every personal-data variable a job
One useful discipline is to require a purpose for each personal-data field before it enters the data collection instrument. That purpose might be analytical, operational, safety-related, regulatory, or methodological.
| Possible variable |
A defensible reason might be |
Question to ask before collecting it |
| Age |
Eligibility depends on age, age is an explanatory variable, or age adjustment is specified in the analysis. |
Do you need exact age or date of birth, or would an age range be sufficient? |
| Email address |
Participants must receive appointments, follow-up surveys, study results, or compensation information. |
Does the email address need to remain attached to the research dataset? |
| Exact home address |
The research genuinely requires household-level location or an intervention must be delivered there. |
Would city, municipality, district, postal area, or another less precise location answer the research question? |
| Occupation |
Occupation is part of the hypothesis, sampling strategy, exposure assessment, or planned subgroup analysis. |
Do you need an exact job title, or would a broader occupational category work? |
| Income |
Socioeconomic position is substantively relevant to the research question or a specified confounder. |
Is exact income necessary, or would meaningful income bands provide sufficient analytical precision? |
| Student or employee number |
Records must be linked across an authorized source or participants must be tracked longitudinally. |
Can a study-generated identifier accomplish the task instead? |
The examples are not rules about what researchers may or may not collect. Their purpose is to expose a more useful question: what level of detail is necessary for the particular function?
Demographic questions are not automatically necessary
Researchers frequently add standard demographic sections almost by reflex. Age, sex, gender, marital status, educational attainment, employment, income, ethnicity, religion, and location may all be important in particular studies. None is automatically necessary simply because demographic tables are common in published papers.
Ask what you intend to do with each variable. Will it describe the sample in a substantively meaningful way? Is it part of a hypothesis? Will it be used in sampling, adjustment, stratification, effect-modification analysis, or interpretation? Is it needed to assess representativeness or inequity? Is there another legitimate requirement?
If the only explanation is "we usually include it," reconsider the variable. Established disciplinary practice can inform study design, but habit alone does not demonstrate necessity.
Collecting enough data is also part of good research
Minimizing personal data does not mean indiscriminately deleting variables until privacy risk approaches zero. Under the GDPR formulation, for example, personal data should be "adequate, relevant and limited to what is necessary" for the purpose. Adeacy matters alongside limitation.
If age is a genuine confounder, refusing to collect any age information in the name of privacy could weaken the study. If longitudinal follow-up is essential, some mechanism for reconnecting records may be necessary. If a study examines disparities between groups, relevant demographic characteristics may be scientifically indispensable.
The goal is therefore not the smallest imaginable dataset. It is the smallest dataset that still properly fulfils the defined purpose.
Ask whether you need the exact value
Sometimes the variable is necessary but its precision is not. A researcher may need participants' ages without needing full dates of birth. Geographic context may matter without requiring exact residential addresses. Socioeconomic status may be analyzable using categories rather than exact financial figures.
This creates an important distinction between needing information about a characteristic and needing its most precise possible value. Reducing precision can sometimes preserve analytical utility while reducing identifiability or sensitivity.
Ask whether you need identifying information in the analytical dataset
A project may legitimately require identifiers at one stage without needing them everywhere. Recruitment staff may need names and contact information. Analysts may need only study IDs and research variables.
Separating contact or identifying information from analytical data can reduce unnecessary exposure. Pseudonymization may also be appropriate in some projects, although pseudonymized information generally remains personal data under GDPR-style frameworks when reidentification remains possible.
This is one reason to think about responsibility for personal research data at the design stage rather than after collection has begun. Different members of the team do not necessarily need access to the same information.
"It might be useful later" is usually not enough
Exploratory research can legitimately require flexibility, and not every future analysis can be predicted with perfect precision. That does not make unlimited collection defensible.
UK Information Commissioner's Office guidance on data minimisation explicitly advises against collecting personal data merely on the off-chance that they might become useful. Its research guidance also explains that necessity must amount to more than something being useful or habitual: the processing should be a targeted and proportionate means of achieving the research purpose.
The Philippine Data Privacy Act and its implementing rules use the related principle of proportionality. Personal data processing should be adequate, relevant, suitable, necessary, and not excessive in relation to a declared and specified purpose. The implementing rules further state that only personal data necessary and compatible with that purpose should be collected.
Different legal frameworks use somewhat different terminology, but both illustrate the same practical discipline for research design: collect information because you can explain why you need it.
Sensitive information deserves an even stronger justification
Health information, genetic or biometric information, racial or ethnic origin, political or religious information, sexual-life information, and other legally protected categories vary across jurisdictions. Their processing may trigger additional requirements.
Even apart from legal classification, some information can create greater harm if disclosed or misused. Researchers should therefore ask whether a less sensitive variable could answer the same question, whether a broader category would suffice, and whether the information needs to remain identifiable.
Watch Out
Do not assume that participant consent makes unnecessary data collection harmless. Consent and necessity answer different questions. Applicable law, ethics requirements, and institutional policies may still limit what should be collected and how it may be processed.
Data requirements can change during the research lifecycle
A field may be necessary during recruitment but unnecessary after eligibility is confirmed. Contact information may be needed during follow-up but not after the final participant communication. A linkage key may be needed until datasets are combined and then become unnecessary for most members of the research team.
For that reason, deciding what to collect is only the first step. Researchers should also determine who needs each type of information, for how long, and at which stage of the project. These questions lead directly to data minimization across the research lifecycle.
04 · A Practical Example
Turning a 12-Field Participant Profile Into What the Study Actually Needs
Hypothetical Example
A study of university students' study habits and academic stress
A researcher drafts a survey asking for full name, university email, student number, exact date of birth, age, sex, gender, home address, degree program, year level, household income, and religion before the main measures of study habits and academic stress.
Define the analysis
The research questions require age group, degree program, year level, study-habit measures, and academic-stress scores. No hypotheses or planned analyses involve religion, exact residential location, household income, sex, or gender.
Separate operational needs
Email addresses are needed only because participants who request a summary of findings will receive one later. They are not required for analysis.
Reduce unnecessary precision
The analysis requires age categories rather than exact birth dates. The survey therefore asks for the required age information without collecting a full date of birth.
Remove unexplained variables
The researcher removes fields that have no defined methodological or operational purpose rather than keeping them for unspecified future analyses.
Separate identifiers
Email addresses for participants requesting results are stored separately from survey responses using an appropriate study process.
This does not mean the removed variables are inherently inappropriate research variables. A study of gender differences in academic stress could have a clear reason to collect gender. A socioeconomic inequality study could require an income measure. The justification changes because the research purpose changes.
06 · What This Means for You
Make Every Personal-Data Field Defend Its Place
Before finalizing a questionnaire, interview protocol, extraction form, database, or data request, review the personal-data variables one by one. The useful question is not simply, "Could this be useful?" but "What defined purpose requires this information at this level of detail?"
A simple decision framework
If the variable directly measures something required by the research question
Collect the level of information necessary for valid measurement, subject to applicable ethical and legal requirements.
If the variable is needed for eligibility, sampling, confounding control, stratification, linkage, follow-up, or safety
Document that purpose and ask whether a less identifying or less precise version would still work.
If the variable is included only because similar studies collect it
Reassess it. Convention can inform your reasoning but should not replace it.
If you cannot explain what you will do with the information
Do not collect it merely in case it becomes interesting later.
Then repeat the exercise for precision. You may need age but not date of birth, region but not street address, an occupational category but not an employer's name. The appropriate choice depends on what the research actually needs.
Finally, separate collection from access. Even where the project legitimately collects identifiable information, that does not mean every collaborator needs access to participant-level identifiers. What the project needs and what each individual team member needs are different questions.
07 · A Quick Checklist
Before Adding a Personal-Data Field
For each personal-data variable, check:
State the specific research, operational, safety, regulatory, or methodological purpose the variable serves.
Confirm that the variable is actually relevant to that purpose rather than merely potentially interesting.
Ask whether the purpose can reasonably be achieved without collecting personal data at all.
Ask whether a broader category, less precise value, or study-generated identifier would be sufficient.
Verify whether the information is legally sensitive or otherwise presents heightened privacy or participant risk.
Determine whether identifiers can be stored separately from the analytical dataset.
Define who actually needs access to the variable and during which stage of the project.
Check applicable ethics requirements, data protection law, and institutional policies before collection begins.