Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

When Should Participant Identifiers Be Separated From Research Data?

Participant identifiers should generally be separated from substantive research data as early as practical when direct identity is no longer needed for the task being performed. The appropriate point depends on whether the study still requires participant-specific contact, linkage, verification, or other authorised functions.

283
When to Separate Participant Identifiers Guide 283 of 398
01 · The Question

At What Point Should Names and Other Identifiers Leave the Working Dataset?

A study may legitimately begin with identifiable information. Researchers recruit named participants, obtain consent, schedule interviews, link records, or follow people over time.

But identity does not necessarily need to accompany the data through every later stage.

Should identifiers be separated immediately after enrolment? After data collection? Before transcription? Before analysis? Only after longitudinal follow-up ends?

There is no single moment that fits every study. The useful principle is to separate direct identifiers once they are no longer necessary for the particular research operation, while preserving controlled linkage when the study still legitimately requires it.

02 · The Short Answer

Separate Identifiers as Early as the Study Can Function Without Them

In Brief

Participant identifiers should generally be separated from substantive research data as early as practical once direct identity is no longer necessary for the task, while any required ability to reconnect records to participants is maintained through an appropriately protected linkage mechanism.

The right time depends on the study. Cross-sectional analysis may permit early separation, while longitudinal follow-up, participant-specific procedures, withdrawal mechanisms, record linkage, or safety requirements may justify retaining controlled re-identification for longer.

03 · What You Need to Know

Separation Should Follow the Research Workflow

Separation Means Breaking the Routine Proximity Between Identity and Research Content

Separating identifiers does not necessarily mean destroying them.

A research team might remove names, email addresses, telephone numbers, medical record numbers, or other direct identifiers from the working dataset and replace them with a study code. A separate, more restricted file then preserves the relationship between the code and the participant when authorised re-identification remains necessary.

This is commonly a form of coding or pseudonymization rather than anonymization.

Under the UK GDPR framework, pseudonymization means processing personal data so that they can no longer be attributed to a particular person without additional information, provided that additional information is kept separately and subject to technical and organisational measures. The data remain personal data.

The Best Time Is Usually Before People Who Do Not Need Identity Begin Working With the Data

One practical trigger for separation is a change in function.

Recruitment staff may need names and contact details. Once survey responses move to the analysis team, those identifiers may serve no analytical purpose. Separating them before routine analysis prevents identity from accompanying the data merely because it was needed earlier.

This principle can be expressed simply: if the next person or process does not need to know who the participant is, consider removing direct identity before the information reaches that person or process.

That supports role-based access to identifiable research information rather than relying only on promises that authorised users will ignore columns they do not need.

Some Identifiers Can Be Separated Immediately After Collection

Consider a one-time questionnaire in which names are collected only to document enrolment or administer incentives. If no later research procedure requires researchers to connect a participant's identity with their responses, direct identifiers may be separable very early.

In some designs, they may not need to be collected with the responses at all.

Researchers should therefore ask during protocol development whether identifying information can follow a different data path from substantive research information from the beginning.

Longitudinal Studies Often Need the Link for Longer

A longitudinal study must usually recognise that the person who completed Wave 1 is the same person completing Wave 2. The research team may also need to send reminders and schedule later assessments.

That does not require every longitudinal dataset to contain names.

A participant can be assigned a stable study code. The analysis dataset uses that code to connect observations across time, while a separate file connects the code to contact information or identity for authorised follow-up.

The identity link remains, so the data are not necessarily anonymous. Separation instead reduces routine exposure while preserving the longitudinal function.

Identifiers Used Only for Contact Can Often Follow a Separate Path

Email addresses and telephone numbers frequently exist for recruitment, scheduling, reminders, or incentives rather than analysis.

When that is the case, researchers can consider storing participant contact information separately from research responses.

The key design question is whether the contact information needs to remain linkable to responses. Follow-up studies may require that connection. An anonymous survey collecting a separate prize-draw email may not.

Transcription Creates Its Own Separation Decision

Qualitative studies often begin with inherently identifiable source material. An interviewer knows whom they interviewed, and an audio recording may contain a recognisable voice, names, workplaces, relationships, locations, or other identifying details.

A research team may decide to create a working transcript in which unnecessary direct identifiers are removed or replaced with pseudonyms before coding and thematic analysis. The original recording and any identity key can then receive more restricted access than the working transcript.

This does not guarantee anonymity. Narrative details can still identify participants without their names appearing. Separation nevertheless reduces unnecessary exposure of obvious identity information during routine qualitative analysis.

Record Linkage May Require Temporary Access to Identifiers

Some research deliberately connects participant information with health records, educational records, administrative databases, registries, or other sources.

Linkage may require identifiable variables during the matching process. Once the authorised linkage has been completed, the analytical dataset may no longer need those direct identifiers.

One possible architecture therefore separates functions:

Identity information Used by authorised personnel or a linkage service to perform the approved match.
Linkage output Records receive a study-specific code or another mechanism needed for longitudinal or cross-dataset analysis.
Analysis dataset Direct identifiers are excluded where they are no longer analytically necessary.

The exact process should follow the approved protocol, applicable law, institutional requirements, and data-provider agreements.

Separation Should Happen Before External Sharing Whenever Direct Identity Is Unnecessary

Collaborators and repositories should not receive direct identifiers merely because the originating research team possesses them.

Before data move to another person or organisation, researchers should determine what the recipient actually needs. Depending on the purpose and governing requirements, sharing may involve coded, pseudonymized, de-identified, anonymized, aggregated, or otherwise appropriately prepared information.

The terminology matters because separating direct identifiers does not necessarily remove re-identification risk from indirect identifiers and linked information.

Watch Out

Do not assume that removing names immediately before sending a dataset makes it safe to share. Exact dates, detailed geography, rare characteristics, free text, metadata, code keys, and combinations of variables may still make participants identifiable. Separation of direct identifiers is one safeguard, not a complete disclosure-risk assessment.

Separate the Key From the Coded Data in a Meaningful Way

If the code key is stored beside the coded dataset with identical access permissions, the files may be technically separate while providing little practical access separation.

Meaningful separation can involve different permissions, storage locations within approved institutional systems, role restrictions, or other technical and organisational controls appropriate to the study.

The aim is to prevent every routine user of the research data from automatically acquiring the information needed to restore identity.

Under OHRP guidance, a code key is precisely what permits coded private information to be linked back to individuals. In certain secondary-research circumstances, whether investigators can obtain that key can determine whether they can readily ascertain identities for purposes of the US Common Rule.

Separation and Permanent Removal Are Different Decisions

A study may separate identifiers early while retaining the linkage key for years because follow-up, withdrawal, safety monitoring, audit requirements, or another approved function remains necessary.

Later, the research team may reach a point at which that connection is no longer needed. Permanently destroying the linkage mechanism is a more consequential step because it may prevent future re-identification by the research team.

Separation Identity information is kept apart from substantive research data, while authorised linkage may remain possible.
Permanent removal The research team eliminates a particular identity linkage when it is no longer needed, subject to the remaining data and applicable requirements.

The question of when identifiers should be permanently removed therefore comes after, and is distinct from, deciding when they should stop accompanying routine research data.

Separation Can Occur at Different Times for Different Identifiers

There is no requirement that every identifier follow the same schedule.

A home address needed only to mail study materials may become unnecessary after enrolment. An email address may remain necessary through follow-up. A participant code may remain throughout analysis. A signed consent record may have a separate retention requirement.

Thinking of "the identifiers" as one indivisible package can therefore be misleading. Each category can have its own purpose and lifecycle.

The Right Timing Depends on What Would Be Lost by Breaking the Link

Separating identifiers from the working dataset is often reversible because the protected key remains. Permanently eliminating the connection is not.

Before changing the architecture, researchers should identify what participant-specific functions remain necessary.

Research Stage Why Identity May Still Be Needed Possible Separation Approach
Recruitment Invitation, eligibility, scheduling Keep contact information in a recruitment system separate from later analytical data where practical
Data collection Participant-specific procedures or follow-up Assign a study code and limit direct identifiers to authorised functions
Transcription Source recordings may contain identity Create a working transcript with unnecessary direct identifiers removed or replaced where appropriate
Longitudinal follow-up Contact and repeated-measures linkage Use stable study codes while protecting the identity key separately
Routine analysis Identity often has no analytical role Provide analysts with coded data without unnecessary direct identifiers
External collaboration Depends on collaborative task Share only the level of identifiability necessary and authorised for that purpose
Project closure Retention, audit, integrity, future approved use, or other obligations may remain Review whether continued linkage is justified before permanent removal
04 · A Practical Example

Separating Identifiers at Different Points in a Longitudinal Study

Hypothetical Example

A Two-Year Study of Doctoral Student Well-Being

Researchers enrol doctoral students and collect surveys every six months. Participants may withdraw during follow-up, so the team needs to retain a controlled connection between identities and research records.

At enrolment Authorised staff collect names and email addresses and assign each participant a random study code.
After each survey Research responses are stored under the study code rather than the participant's name or email address.
During follow-up A restricted contact file allows authorised staff to send the next invitation and, where applicable, locate records for participant-specific requests.
During analysis Analysts receive longitudinal records linked by study code. They do not need the contact file to model changes across the four survey waves.
After participant-specific functions end The team reviews whether continued linkage is required under the protocol, retention plan, institutional policy, applicable law, or future approved research arrangements before deciding whether identifiers should be permanently removed.

The identifiers were separated early even though the identity link was retained. Separation reduced routine exposure without sacrificing functions the longitudinal design still required.

05 · What Researchers Often Get Wrong

Common Mistakes About Separating Participant Identifiers

Misconception

Identifiers Should Stay With the Dataset Until the Study Ends

Not necessarily. If direct identity is needed for recruitment but irrelevant to analysis, identifiers may be separated much earlier while authorised linkage is preserved separately when necessary.

Misconception

Identifiers Must Be Permanently Deleted as Soon as They Are Separated

Separation and deletion solve different problems. A study may need continued re-identification for longitudinal follow-up, withdrawal, safety procedures, linkage, or other approved functions while still keeping identifiers out of routine working data.

Misconception

A Participant Code Makes Separation Unnecessary

A study code is useful precisely because it can allow research data to function without direct identifiers. If the code key sits in the same unrestricted dataset, much of that access-control benefit may be lost.

Misconception

Removing Names Before Analysis Makes the Data Anonymous

Not automatically. A retained key, indirect identifiers, distinctive narratives, metadata, or other information may still permit identification. Separation is usually a confidentiality or pseudonymization measure rather than proof of anonymity.

Misconception

All Identifiers Should Be Separated at the Same Time

Different identifiers serve different purposes. Contact details, addresses, linkage codes, consent documentation, and other identifying information may legitimately have different access and retention lifecycles.

06 · What This Means for You

Use the Next Research Task to Decide Whether Identity Still Belongs With the Data

Instead of choosing one arbitrary project date for "de-identification," examine each transition in the data workflow.

A simple decision framework

If direct identity is needed only for recruitment or scheduling
Consider separating it before substantive research data move into routine processing and analysis.
If the study requires repeated follow-up or participant-specific actions
Preserve an appropriately protected linkage mechanism while keeping direct identifiers out of datasets that do not need them.
If qualitative source material contains unnecessary direct identifiers
Consider creating a working version with those identifiers removed or replaced before wider coding and analysis, while preserving source material according to the approved protocol.
If an external recipient does not need participant identities
Prepare the data to exclude unnecessary direct identifiers and separately assess remaining indirect identification risk before sharing.
If no authorised function requires the identity link anymore
Evaluate whether permanent removal is appropriate under the study's retention, integrity, legal, ethical, and institutional requirements.

The practical rule is straightforward: direct identity should not accompany a research operation merely because an earlier operation needed it. Treat identity as a functional requirement that must continue to justify its presence at each stage.

07 · A Quick Checklist

Identify the Earliest Practical Point for Separation

For each stage of the research workflow, check:
Identify which direct identifiers are still necessary for the next research or operational task.
Assign study codes when authorised linkage is necessary but routine use of direct identity is not.
Store the code key or other linkage information separately with access appropriate to its function.
Remove unnecessary direct identifiers before routine analysis where the analytical task does not require them.
Review qualitative transcripts and other working materials for direct identifiers before broader team access where appropriate.
Before external sharing, exclude identifiers the recipient does not need and assess remaining indirect identification risk.
Give different identifiers separate lifecycles rather than assuming all identifying information must follow the same schedule.
Document when and why separation occurs in the data-management plan or other study documentation where appropriate.
Verify the applicable ethics, legal, institutional, sponsor, repository, and retention requirements before permanently breaking an identity link.
08 · Frequently Asked Questions

Frequently Asked Questions About Separating Participant Identifiers

Should identifiers be separated immediately after data collection?

Often they can be, but not universally. If direct identity is no longer needed for the next research task, early separation can reduce exposure. Studies requiring follow-up, linkage, safety procedures, or other participant-specific functions may need to preserve controlled re-identification.

Can identifiers be separated before data collection?

In some designs, yes. Contact or administrative information can sometimes follow a separate collection path from substantive research responses from the outset. Whether this is appropriate depends on whether the study needs to connect the two.

Should identifiers be removed before statistical analysis?

Direct identifiers often have no analytical purpose and can frequently be excluded from routine analysis datasets. Study codes can preserve necessary longitudinal or cross-dataset linkage. Some specialised analyses may require different arrangements.

Should names be removed from interview transcripts?

Removing or replacing unnecessary direct identifiers before wider analysis can reduce exposure, but qualitative narratives may remain identifiable through contextual details. The original recording and transcript may also have distinct retention and access requirements.

Does separating identifiers make research data anonymous?

No, not when a linkage key remains or the data permit identification through other means. Separation commonly produces coded or pseudonymized information rather than anonymous information.

Where should the code key be stored?

It should be stored using institutionally approved arrangements appropriate to its sensitivity and re-identification function, with access limited to authorised people who need it. Specific technical requirements depend on the institution, study, and applicable rules.

When should the code key be destroyed?

Only when the identity link is no longer needed and permanent removal is consistent with the approved protocol, participant commitments, retention requirements, research-integrity needs, applicable law, and institutional policy. Separation can occur much earlier than destruction of the key.

09 · The Bottom Line

Separate Identity When the Next Task No Longer Needs It

The Bottom Line

Participant identifiers should generally be separated from substantive research data as soon as direct identity is no longer necessary for the research operation being performed, while authorised linkage can be preserved separately when the study still requires it.

Think in stages rather than waiting for one universal "de-identification date." Recruitment, data collection, transcription, follow-up, analysis, sharing, and project closure may each justify a different relationship between identity and research data.

10 · Sources and Further Reading

Authoritative Sources on Separating Identifiers and Research Data

Is this conversation helpful so far?
11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes