03 · What You Need to Know
Data Minimization Is a Purpose Test, Not a Data-Count Test
The principle begins with a defined purpose
You cannot determine whether data are excessive until you know what they are supposed to accomplish. A dataset that is excessive for one study may be entirely justified for another.
Under the GDPR formulation, personal data must be "adequate, relevant and limited to what is necessary" in relation to the purposes for which they are processed. European Commission guidance summarizes this as collecting and processing only the personal data necessary to fulfil the purpose.
The Philippine framework uses somewhat different language but reaches a related practical requirement through proportionality. The implementing rules of the Data Privacy Act state that processing should be adequate, relevant, suitable, necessary, and not excessive in relation to a declared and specified purpose. They further provide that only personal data necessary and compatible with that purpose should be collected.
These are jurisdiction-specific legal formulations, not a universal definition applicable everywhere. Researchers should identify the law governing their project. As a research-design principle, however, they illustrate why minimization begins by asking what the processing is actually for.
"Adequate" prevents minimization from becoming under-collection
The first word in the GDPR formulation is easy to overlook. Data must be adequate. The information must be sufficient to fulfil the stated purpose properly.
Suppose a study examines whether an intervention works differently across age groups. Age may be necessary. Removing it simply because it is personal data could make the planned analysis impossible or scientifically weaker. Minimization does not require that result.
Conversely, the same study may need age without needing a participant's full date of birth. The scientific requirement establishes what information must be available; minimization asks whether additional precision contributes enough to the purpose to justify processing it.
Too little data
The study cannot properly fulfil its defined research purpose.
Minimized data
The study has sufficient relevant information for the purpose without unnecessary personal data or detail.
"Relevant" requires a rational connection to the purpose
A variable can be easy to collect and academically interesting while still having no meaningful relationship to the stated purpose. Researchers should be able to explain that relationship.
If religion, marital status, ethnicity, income, political views, exact location, or another personal characteristic is relevant to the research question, sampling design, confounding structure, equity analysis, eligibility criteria, or another defined function, its collection may be justifiable subject to applicable requirements. If it is included merely because demographic questions usually appear at the end of a survey, its relevance is much harder to demonstrate.
This is where deciding which personal-data variables the study actually needs becomes a methodological as well as a privacy decision.
"Limited to what is necessary" asks whether you can reasonably do with less
Necessity should not be reduced to whether information could conceivably be useful. ICO research guidance explains that, in contexts where necessity is required, processing must be more than useful or habitual. It should be a targeted and proportionate way of achieving the purpose, and researchers should consider whether the same purpose could reasonably be achieved by less intrusive means.
That can lead to several kinds of minimization. You might remove an unnecessary variable, reduce its precision, replace a direct identifier with a study code, separate identifiers from analytical data, limit access to particular team members, or disclose only the fields a collaborator needs.
Minimization is not only about the questionnaire
The most obvious application is deciding what to collect, but the principle concerns processing more broadly. Personal data can become excessive at later stages even when their original collection was justified.
| Research stage |
Minimization question |
Possible response |
| Collection |
Do we need this variable? |
Remove fields with no defined purpose. |
| Measurement |
Do we need this level of precision? |
Use age bands instead of full birth dates when appropriate. |
| Identification |
Does this stage require direct identifiers? |
Use study IDs or pseudonymization where suitable. |
| Access |
Does every team member need every field? |
Use role-based or otherwise appropriately restricted access. |
| Sharing |
Does the recipient need the entire dataset? |
Share only the variables and records required for the agreed purpose. |
| Retention |
Do we still need identifiable information? |
Delete, anonymize, or otherwise reduce data when the applicable purpose and retention requirements permit. |
Data minimization is therefore better understood as a continuing design constraint than as a one-time questionnaire cleanup.
Minimization can reduce the consequences of a breach
Every additional personal-data field creates another piece of information that may be exposed, misused, incorrectly shared, or accessed by someone who does not need it. The severity of that risk depends on the information and context, not simply the number of columns.
Imagine two stolen research files. One contains study IDs, age groups, and research scores. The other contains the same information plus names, exact dates of birth, home addresses, telephone numbers, personal email addresses, and several sensitive characteristics that the analysis never required. The security failure is serious in either case, but unnecessary information can enlarge the consequences of the incident.
Minimization therefore complements security rather than replacing it. You still need appropriate protection for the data you legitimately hold. Collecting less unnecessary information simply means there is less unnecessary information available to be compromised.
Anonymization and pseudonymization can support minimization, but they are different
Where research can be performed using genuinely anonymous information, personal-data obligations under GDPR-style frameworks may no longer apply to that anonymous dataset. Genuine anonymization, however, requires more than deleting names.
Pseudonymization replaces or separates identifying information so that data cannot be attributed to a person without additional information. It can substantially reduce risk and is explicitly recognized in research safeguards under GDPR-related provisions, but pseudonymized data remain personal data when reidentification remains possible.
The appropriate approach depends on the research. Longitudinal follow-up or record linkage may require a controlled ability to reconnect records to individuals. Other analyses may not need that capability at all.
Research provisions do not erase the minimization principle
Some data protection regimes provide particular provisions, safeguards, or exceptions for scientific, historical, statistical, or public-interest research. These should not be interpreted as permission to accumulate unlimited personal data.
For example, UK research guidance identifies measures respecting data minimization, including anonymization and pseudonymization where possible, as important safeguards for research-related processing. The Philippine implementing rules likewise tie research-related processing to safeguards, declared purposes, confidentiality, and proportionality in relevant circumstances.
The exact rules vary, so researchers should verify their own jurisdiction rather than borrowing an exemption from another country's law.
Minimization also applies when data move to someone else
A research team may legitimately hold a rich dataset while a statistician needs only selected variables, a transcription service needs only audio files, or a collaborator needs only de-identified measurements for a particular analysis.
Membership in the research team does not automatically establish a need for unrestricted access. Before giving another researcher participant-level information, determine which participant data that collaborator actually needs and what governance applies to the disclosure.
The same reasoning matters when transferring research data between institutions. Minimization can sometimes reduce the amount or sensitivity of information that must cross organizational boundaries.
04 · A Practical Example
Minimizing a Longitudinal Research Dataset Without Losing Follow-Up
Hypothetical Example
A six-month follow-up study
A research team surveys 800 participants at baseline and again six months later. The team needs to match each participant's two responses and send a follow-up invitation. The initial database design places names, email addresses, telephone numbers, full dates of birth, home addresses, survey responses, and study outcomes in one spreadsheet.
Purpose
The researchers determine that names and email addresses are needed for follow-up, while a study ID is sufficient to match research records. Telephone numbers and home addresses have no defined function.
Collection
Unnecessary contact fields are removed. Date-of-birth information is reduced to the age information genuinely required by the analysis.
Separation
The contact list linking participant identities to study IDs is stored separately from the analytical dataset with appropriate access restrictions.
Access
Staff responsible for participant follow-up can access the contact information they require. Analysts work with study IDs and research variables rather than direct identifiers.
Later review
Once follow-up and any other justified identifying functions are complete, the team follows its approved retention plan and applicable requirements to determine whether the identifying link should be deleted, retained securely, or otherwise handled.
The study still collects personal data because its design requires them. Minimization does not prevent longitudinal research. It asks the researchers to separate what is scientifically or operationally necessary from what was simply convenient to place in the same file.
06 · What This Means for You
Apply Minimization Every Time the Data Change Hands or Purpose
For practical research management, minimization can be turned into a recurring question: what is the least personal information that allows this particular task to be done properly?
A simple decision framework
If the research can achieve its purpose without personal data
Consider using appropriately anonymized or aggregate information instead.
If personal data are necessary
Identify the specific variables and level of detail required rather than collecting broadly.
If identifiers are needed only temporarily or for one function
Consider separating them from research data and restricting access appropriately.
If another person or institution needs data
Determine whether a reduced dataset can accomplish the agreed purpose before transferring the full dataset.
If information no longer serves its justified purpose
Apply the approved retention and disposal arrangements and any applicable legal, ethical, institutional, or archival requirements.
Documenting these decisions can be valuable. If you decide that exact age, detailed location, or a particular sensitive characteristic is necessary, record why. Minimization does not require researchers to apologize for scientifically necessary data. It requires them to be able to distinguish necessary processing from accumulation by default.
That discipline becomes particularly important when selecting storage systems or sharing arrangements. Whether data can be placed in a commercial cloud service is a separate question from whether every field needs to be uploaded there in the first place.
07 · A Quick Checklist
How to Apply Data Minimization to a Research Project
Before and during personal-data processing, check:
Define the specific purpose before deciding which personal data are needed.
Ask whether the research can reasonably achieve that purpose without personal data.
Confirm that every personal-data variable has a rational connection to a defined purpose.
Check whether less precise, less identifying, or less sensitive information would be sufficient.
Separate direct identifiers from analytical data where appropriate and feasible.
Restrict access according to what each role actually needs.
When sharing data, determine whether a reduced dataset can fulfil the recipient's purpose.
Review personal data during the project and follow applicable retention or deletion requirements when information is no longer needed.
Document why unusually detailed, identifying, or sensitive information is necessary when the justification may not be obvious.