03 · What You Need to Know
Separation Should Follow the Research Workflow
Separation Means Breaking the Routine Proximity Between Identity and Research Content
Separating identifiers does not necessarily mean destroying them.
A research team might remove names, email addresses, telephone numbers, medical record numbers, or other direct identifiers from the working dataset and replace them with a study code. A separate, more restricted file then preserves the relationship between the code and the participant when authorised re-identification remains necessary.
This is commonly a form of coding or pseudonymization rather than anonymization.
Under the UK GDPR framework, pseudonymization means processing personal data so that they can no longer be attributed to a particular person without additional information, provided that additional information is kept separately and subject to technical and organisational measures. The data remain personal data.
The Best Time Is Usually Before People Who Do Not Need Identity Begin Working With the Data
One practical trigger for separation is a change in function.
Recruitment staff may need names and contact details. Once survey responses move to the analysis team, those identifiers may serve no analytical purpose. Separating them before routine analysis prevents identity from accompanying the data merely because it was needed earlier.
This principle can be expressed simply: if the next person or process does not need to know who the participant is, consider removing direct identity before the information reaches that person or process.
That supports role-based access to identifiable research information rather than relying only on promises that authorised users will ignore columns they do not need.
Some Identifiers Can Be Separated Immediately After Collection
Consider a one-time questionnaire in which names are collected only to document enrolment or administer incentives. If no later research procedure requires researchers to connect a participant's identity with their responses, direct identifiers may be separable very early.
In some designs, they may not need to be collected with the responses at all.
Researchers should therefore ask during protocol development whether identifying information can follow a different data path from substantive research information from the beginning.
Longitudinal Studies Often Need the Link for Longer
A longitudinal study must usually recognise that the person who completed Wave 1 is the same person completing Wave 2. The research team may also need to send reminders and schedule later assessments.
That does not require every longitudinal dataset to contain names.
A participant can be assigned a stable study code. The analysis dataset uses that code to connect observations across time, while a separate file connects the code to contact information or identity for authorised follow-up.
The identity link remains, so the data are not necessarily anonymous. Separation instead reduces routine exposure while preserving the longitudinal function.
Identifiers Used Only for Contact Can Often Follow a Separate Path
Email addresses and telephone numbers frequently exist for recruitment, scheduling, reminders, or incentives rather than analysis.
When that is the case, researchers can consider storing participant contact information separately from research responses.
The key design question is whether the contact information needs to remain linkable to responses. Follow-up studies may require that connection. An anonymous survey collecting a separate prize-draw email may not.
Transcription Creates Its Own Separation Decision
Qualitative studies often begin with inherently identifiable source material. An interviewer knows whom they interviewed, and an audio recording may contain a recognisable voice, names, workplaces, relationships, locations, or other identifying details.
A research team may decide to create a working transcript in which unnecessary direct identifiers are removed or replaced with pseudonyms before coding and thematic analysis. The original recording and any identity key can then receive more restricted access than the working transcript.
This does not guarantee anonymity. Narrative details can still identify participants without their names appearing. Separation nevertheless reduces unnecessary exposure of obvious identity information during routine qualitative analysis.
Record Linkage May Require Temporary Access to Identifiers
Some research deliberately connects participant information with health records, educational records, administrative databases, registries, or other sources.
Linkage may require identifiable variables during the matching process. Once the authorised linkage has been completed, the analytical dataset may no longer need those direct identifiers.
One possible architecture therefore separates functions:
Identity information
Used by authorised personnel or a linkage service to perform the approved match.
Linkage output
Records receive a study-specific code or another mechanism needed for longitudinal or cross-dataset analysis.
Analysis dataset
Direct identifiers are excluded where they are no longer analytically necessary.
The exact process should follow the approved protocol, applicable law, institutional requirements, and data-provider agreements.
Separation Should Happen Before External Sharing Whenever Direct Identity Is Unnecessary
Collaborators and repositories should not receive direct identifiers merely because the originating research team possesses them.
Before data move to another person or organisation, researchers should determine what the recipient actually needs. Depending on the purpose and governing requirements, sharing may involve coded, pseudonymized, de-identified, anonymized, aggregated, or otherwise appropriately prepared information.
The terminology matters because separating direct identifiers does not necessarily remove re-identification risk from indirect identifiers and linked information.
Watch Out
Do not assume that removing names immediately before sending a dataset makes it safe to share. Exact dates, detailed geography, rare characteristics, free text, metadata, code keys, and combinations of variables may still make participants identifiable. Separation of direct identifiers is one safeguard, not a complete disclosure-risk assessment.
Separate the Key From the Coded Data in a Meaningful Way
If the code key is stored beside the coded dataset with identical access permissions, the files may be technically separate while providing little practical access separation.
Meaningful separation can involve different permissions, storage locations within approved institutional systems, role restrictions, or other technical and organisational controls appropriate to the study.
The aim is to prevent every routine user of the research data from automatically acquiring the information needed to restore identity.
Under OHRP guidance, a code key is precisely what permits coded private information to be linked back to individuals. In certain secondary-research circumstances, whether investigators can obtain that key can determine whether they can readily ascertain identities for purposes of the US Common Rule.
Separation and Permanent Removal Are Different Decisions
A study may separate identifiers early while retaining the linkage key for years because follow-up, withdrawal, safety monitoring, audit requirements, or another approved function remains necessary.
Later, the research team may reach a point at which that connection is no longer needed. Permanently destroying the linkage mechanism is a more consequential step because it may prevent future re-identification by the research team.
Separation
Identity information is kept apart from substantive research data, while authorised linkage may remain possible.
Permanent removal
The research team eliminates a particular identity linkage when it is no longer needed, subject to the remaining data and applicable requirements.
The question of when identifiers should be permanently removed therefore comes after, and is distinct from, deciding when they should stop accompanying routine research data.
Separation Can Occur at Different Times for Different Identifiers
There is no requirement that every identifier follow the same schedule.
A home address needed only to mail study materials may become unnecessary after enrolment. An email address may remain necessary through follow-up. A participant code may remain throughout analysis. A signed consent record may have a separate retention requirement.
Thinking of "the identifiers" as one indivisible package can therefore be misleading. Each category can have its own purpose and lifecycle.
The Right Timing Depends on What Would Be Lost by Breaking the Link
Separating identifiers from the working dataset is often reversible because the protected key remains. Permanently eliminating the connection is not.
Before changing the architecture, researchers should identify what participant-specific functions remain necessary.
| Research Stage |
Why Identity May Still Be Needed |
Possible Separation Approach |
| Recruitment |
Invitation, eligibility, scheduling |
Keep contact information in a recruitment system separate from later analytical data where practical |
| Data collection |
Participant-specific procedures or follow-up |
Assign a study code and limit direct identifiers to authorised functions |
| Transcription |
Source recordings may contain identity |
Create a working transcript with unnecessary direct identifiers removed or replaced where appropriate |
| Longitudinal follow-up |
Contact and repeated-measures linkage |
Use stable study codes while protecting the identity key separately |
| Routine analysis |
Identity often has no analytical role |
Provide analysts with coded data without unnecessary direct identifiers |
| External collaboration |
Depends on collaborative task |
Share only the level of identifiability necessary and authorised for that purpose |
| Project closure |
Retention, audit, integrity, future approved use, or other obligations may remain |
Review whether continued linkage is justified before permanent removal |