Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Should Participant Contact Information Be Stored With Research Responses?

Participant contact information often serves an administrative purpose rather than an analytical one. When responses do not need to contain names, emails, or telephone numbers, separating contact details can reduce unnecessary exposure while preserving authorised linkage when the study requires it.

281
Storing Contact Information With Research Responses Guide 281 of 398
01 · The Question

Should Names, Emails, and Phone Numbers Sit Beside Participant Responses?

A researcher needs participants' email addresses to schedule follow-up interviews. The survey platform conveniently exports those addresses in the same spreadsheet as every response.

That arrangement works. But does the analysis dataset actually need the email addresses?

Contact information and research responses often serve different purposes. Keeping them together can create an unnecessary direct link between identity and potentially sensitive research information. Separating them can reduce exposure, although some studies legitimately require controlled linkage between the two.

02 · The Short Answer

Usually Separate Them Unless Keeping Them Together Serves a Necessary Purpose

In Brief

Participant contact information generally should not be stored directly with substantive research responses when the study can achieve its purposes by keeping the two separately and linking them only when necessary.

Separation does not make the research anonymous if a code or other mechanism can reconnect the records. It is a confidentiality and data-minimisation safeguard that reduces unnecessary exposure of direct identifiers while preserving authorised linkage where the research requires it.

03 · What You Need to Know

Contact Information and Research Responses Usually Do Different Jobs

Contact Details Are Often Administrative Rather Than Analytical Data

Names, email addresses, telephone numbers, mailing addresses, and similar information may be necessary to recruit participants, schedule appointments, send reminders, deliver incentives, arrange follow-up, or communicate study information.

None of those purposes necessarily requires the contact information to appear in the dataset used for substantive analysis.

Separating administrative information from research responses follows a simple design principle: give each category of information only to the people and systems that need it for its actual purpose.

This also supports the broader principle of collecting only identifying information that the study genuinely needs.

Keeping Contact Information With Responses Creates a Direct Identity Link

Suppose a spreadsheet contains these columns:

  • participant name;
  • email address;
  • telephone number;
  • mental-health score;
  • workplace satisfaction;
  • income;
  • political attitude; and
  • open-ended comments.

Anyone with access to that file immediately receives both identity and research content.

If the same project instead uses a study code in the response dataset and stores contact information separately, routine analysts may be able to work without seeing direct identifiers. The identity connection can be restricted to authorised people who need it for legitimate study functions.

The underlying study may remain identifiable or pseudonymized, but the exposure pathways have changed.

Separation Is Not the Same as Anonymity

This distinction is crucial.

Imagine the research-response file contains participant code P084. A second file contains:

P084 → participant@example.com

The files are separate, but the link remains. Someone with access to both can determine which responses belong to that participant.

Under frameworks such as the UK GDPR, this arrangement is generally closer to pseudonymization than anonymization. Pseudonymized data remain personal data where attribution is possible using separately held additional information.

Separated or pseudonymized Contact information is stored apart from responses, but authorised linkage remains possible.
Anonymous No relevant party can reconnect the research information to an identifiable person under the applicable standard.

Separating contact information is therefore valuable without needing to oversell what it accomplishes.

A Simple Three-File Architecture Can Reduce Unnecessary Access

Where a study needs participant contact and longitudinal linkage, one possible architecture is:

File Contents Who May Need It?
Contact file Study code, name, email, telephone number, or other necessary contact details Authorised staff responsible for participant communication
Research dataset Study code and substantive research variables Researchers and analysts who need the study data
Linkage information Information necessary to connect identifiers and study records, if not already contained in the contact file A restricted subset of authorised personnel

The exact architecture should match the study, applicable law, institutional systems, and security requirements. The point is functional separation: not every person who analyses responses needs access to every identifier.

Sometimes Contact Information Does Not Need to Be Linkable to Responses at All

Consider an anonymous online survey offering participants entry into a prize draw. The researcher needs email addresses to select and contact winners, but the email addresses do not need to be connected to survey answers.

One possible design is to collect survey responses first and direct interested participants to a separate contact form that does not carry a participant identifier or other mechanism connecting the contact submission to the survey response.

If the systems are genuinely separated and no other identifying information is collected, the contact list can exist without necessarily compromising the anonymity of the survey responses.

Watch Out

A second form is not automatically separate merely because it appears on another webpage. Unique URLs, embedded participant IDs, timestamps, account information, platform logs, IP addresses, or other metadata may preserve a connection. Check the actual technical data flow before promising that contact details cannot be linked to responses.

Other Studies Need the Link and Should Preserve It Deliberately

Longitudinal research may need to connect the same person's responses across multiple waves. Clinical or intervention studies may need participant-specific follow-up. Researchers may need to respond to withdrawal requests, return individual results where appropriate, verify eligibility, or connect research data with other authorised records.

In these circumstances, destroying every identity link at collection would interfere with legitimate study functions.

The solution is not to call the data anonymous. It is to retain the necessary link in a controlled manner, explain it accurately to participants, and determine who actually needs access to identifiable information.

Separation Reduces the Consequences of Some Access Events

Suppose an analyst accidentally receives the research dataset. If names and contact details are embedded in every row, the analyst immediately receives direct identifiers as well.

If those identifiers are held separately, obtaining one file may not provide the same direct identity connection. The effectiveness of this safeguard depends on what remains in the research dataset, because indirect identifiers can still make participants identifiable.

Separation therefore reduces one category of risk rather than guaranteeing non-identifiability.

Contact Information Should Have Its Own Access Rules

Researchers sometimes secure "the dataset" while leaving participant contact lists in ordinary email inboxes, spreadsheets, shared drives, calendars, or scheduling systems.

Contact information is itself personal information and should be included in the study's data-management plan. Relevant questions include:

  • where it will be collected;
  • where it will be stored;
  • who can access it;
  • how it will be transferred;
  • whether it is included in backups;
  • how long it is needed; and
  • what happens when its purpose ends.

Separating identifiers from responses accomplishes little if the identifier file is then broadly accessible.

The Linkage Key Deserves Stronger Access Control Than Routine Analysis Data

A linkage file can turn coded research data back into participant-specific information. It should therefore not be treated as an administrative afterthought.

OHRP defines coded private information as information in which identifying information has been replaced by a code while a key exists that enables linkage. Its guidance also illustrates how access to that key can affect whether investigators can readily ascertain participants' identities under the US Common Rule.

The legal consequences vary by framework, but the operational lesson is straightforward: the ability to reconnect identity is itself sensitive research infrastructure.

Separation Can Also Support Role-Based Access

Different research personnel often need different information.

A recruitment coordinator may need participant names and contact information but not complete analytical variables. A statistician may need detailed response data but not names. A principal investigator may require authorised access to both in particular circumstances.

Separating files makes these distinctions technically possible. If everything sits in one spreadsheet, access control becomes binary: either someone gets the whole file or they do not.

This is one reason pseudonymization is recognised in data-protection guidance as an important technical and organisational safeguard even though it does not itself make personal data anonymous.

Eventually, the Identity Link May No Longer Be Necessary

Some studies need contact information during recruitment and follow-up but not indefinitely.

Once participant communication ends, the team should determine whether contact information or the linkage mechanism still serves an approved research, regulatory, integrity, archival, or other legitimate purpose. The answer depends on applicable retention requirements and the study design.

Researchers should not assume either that identifiers must be destroyed immediately or that they should be retained forever. The appropriate question is when participant identifiers can and should be permanently removed under the governing requirements.

04 · A Practical Example

Separating Contact Details Without Losing Longitudinal Follow-Up

Hypothetical Example

A Three-Wave Survey of Early-Career Researchers

A study surveys participants at baseline, six months, and twelve months. Researchers need email addresses to send later survey invitations and must connect each participant's responses across all three waves.

Assign a study code Each participant receives a randomly generated study identifier, such as R4821.
Store contact information separately A restricted contact file connects R4821 with the participant's name and email address.
Use the code in research data The response dataset records R4821 with the participant's survey variables but does not contain the person's name or email address.
Limit the identity function Staff responsible for follow-up can use the protected contact information. Analysts who do not need participant identities receive only the coded research data.
Review the link later After follow-up and other participant-specific functions are complete, the team applies its approved retention plan and determines whether continued identity linkage remains necessary.

The study is not anonymous while authorised re-identification remains possible. The benefit of separation is narrower and more practical: identity is exposed only where the study needs it rather than accompanying every use of the research responses.

05 · What Researchers Often Get Wrong

Common Mistakes When Storing Participant Contact Information

Misconception

Separating Names From Responses Makes the Study Anonymous

Not if a code key or other mechanism can reconnect the files. Separation is an important confidentiality safeguard, but the information may remain pseudonymized or otherwise identifiable.

Misconception

Everyone on the Research Team Needs the Contact List

Access should reflect role and purpose. Analysts may need research variables without participant contact details, while recruitment staff may need contact details without unrestricted access to sensitive responses.

Misconception

A Separate Spreadsheet Is Automatically Secure Separation

Two files stored in the same unrestricted folder and available to the same people may provide little meaningful access separation. Technical and organisational controls should match the intended restriction.

Misconception

Contact Details Are Harmless Because They Are Not Research Outcomes

Names, email addresses, telephone numbers, and addresses are identifying information. Their administrative purpose does not remove the need for appropriate access, security, transparency, and retention controls.

Misconception

Contact Information Must Always Be Destroyed as Soon as Data Collection Ends

Not universally. Some studies may need continued contact, withdrawal mechanisms, follow-up, audit trails, or retention required by applicable rules. The appropriate retention period should follow the study's legitimate purposes and governing requirements rather than a universal timetable.

06 · What This Means for You

Design Contact and Response Data as Separate Functions

Before collecting participant contact information, decide whether the study needs to connect it to research responses and, if so, for how long.

A simple decision framework

If contact information is unnecessary
Do not collect it merely because the survey platform offers fields for it.
If contact information is needed but never needs to be linked to responses
Consider a genuinely separate collection process without a shared participant identifier or other hidden linkage mechanism.
If contact information must be linked to responses
Consider a coded architecture that stores direct identifiers separately and restricts access to the linkage mechanism.
If only some team members need participant identities
Use role-based access so routine analysis does not unnecessarily expose contact information.
If the need for participant contact or linkage has ended
Review whether continued retention is justified under the study's approved plan and applicable requirements.

The correct design depends on the research. The important step is to make the linkage intentional. Contact information should not sit beside sensitive responses simply because one software export happened to put them there.

07 · A Quick Checklist

Before Storing Contact Information With Research Responses

Check whether the identity link is actually necessary:
State why names, email addresses, telephone numbers, or other contact details are required.
Determine whether contact information needs to be linked to substantive responses at all.
If linkage is required, consider using a study code rather than embedding direct identifiers in the analysis dataset.
Store linkage information separately with access appropriate to its re-identification function.
Limit contact information and research-data access according to staff roles and legitimate study functions.
Check whether survey platforms, scheduling systems, or other services create hidden links through metadata, account information, unique URLs, or timestamps.
Describe the actual identity-linkage arrangement accurately in ethics and participant materials where required.
Plan how long contact information and linkage mechanisms need to remain available.
Verify institutional security requirements and applicable data-protection, research-ethics, and retention rules.
08 · Frequently Asked Questions

Frequently Asked Questions About Contact Information and Research Data

Should participant email addresses be stored in the research dataset?

Only when their presence there serves a justified purpose. If email addresses are needed solely for communication, storing them separately from substantive responses can reduce unnecessary exposure of direct identifiers.

Does storing contact information separately make the data anonymous?

Not if a study code, key, or other mechanism can reconnect the files. The arrangement may instead be coded or pseudonymized, with separation functioning as a confidentiality safeguard.

Can I collect emails separately for a prize draw after an anonymous survey?

Potentially. The collection should be genuinely separate if the survey is described as anonymous. Check for participant IDs, unique links, timestamps, account data, IP addresses, or other technical information that could connect the contact form to survey responses.

Can the principal investigator keep the linkage key?

A study may legitimately designate authorised personnel to hold linkage information when re-identification is necessary. Who should hold it, how it should be protected, and how long it should exist depend on the protocol, institutional requirements, and applicable law.

Should consent forms be stored with research responses?

Where signed consent documentation directly identifies participants, storing it separately from substantive research data can reduce unnecessary identity exposure. Specific documentation and retention requirements vary by jurisdiction, institution, sponsor, and study type.

Can analysts receive coded data while another team member keeps contact information?

Yes. This is a common role-based arrangement when analysts do not need identities but authorised staff need to maintain participant contact or linkage. The coded data are not necessarily anonymous merely because the analysts cannot see the contact file.

When should contact information be deleted?

When it no longer serves an approved purpose and no applicable legal, regulatory, institutional, integrity, or archival requirement justifies continued retention. There is no universal deletion date for all research.

09 · The Bottom Line

Identity Should Accompany Research Responses Only When It Needs To

The Bottom Line

Participant contact information generally should be separated from substantive research responses when the study can function without keeping direct identifiers in the same working dataset.

If linkage is necessary, preserve it deliberately through an appropriately protected mechanism and limit access to those who need it. Separation does not create anonymity by itself, but it can prevent every ordinary use of the research data from becoming an ordinary use of the participant's identity.

10 · Sources and Further Reading

Authoritative Sources on Separating Identifiers and Research Data

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes