03 · What You Need to Know
Contact Information and Research Responses Usually Do Different Jobs
Contact Details Are Often Administrative Rather Than Analytical Data
Names, email addresses, telephone numbers, mailing addresses, and similar information may be necessary to recruit participants, schedule appointments, send reminders, deliver incentives, arrange follow-up, or communicate study information.
None of those purposes necessarily requires the contact information to appear in the dataset used for substantive analysis.
Separating administrative information from research responses follows a simple design principle: give each category of information only to the people and systems that need it for its actual purpose.
This also supports the broader principle of collecting only identifying information that the study genuinely needs.
Keeping Contact Information With Responses Creates a Direct Identity Link
Suppose a spreadsheet contains these columns:
- participant name;
- email address;
- telephone number;
- mental-health score;
- workplace satisfaction;
- income;
- political attitude; and
- open-ended comments.
Anyone with access to that file immediately receives both identity and research content.
If the same project instead uses a study code in the response dataset and stores contact information separately, routine analysts may be able to work without seeing direct identifiers. The identity connection can be restricted to authorised people who need it for legitimate study functions.
The underlying study may remain identifiable or pseudonymized, but the exposure pathways have changed.
Separation Is Not the Same as Anonymity
This distinction is crucial.
Imagine the research-response file contains participant code P084. A second file contains:
P084 → participant@example.com
The files are separate, but the link remains. Someone with access to both can determine which responses belong to that participant.
Under frameworks such as the UK GDPR, this arrangement is generally closer to pseudonymization than anonymization. Pseudonymized data remain personal data where attribution is possible using separately held additional information.
Separated or pseudonymized
Contact information is stored apart from responses, but authorised linkage remains possible.
Anonymous
No relevant party can reconnect the research information to an identifiable person under the applicable standard.
Separating contact information is therefore valuable without needing to oversell what it accomplishes.
A Simple Three-File Architecture Can Reduce Unnecessary Access
Where a study needs participant contact and longitudinal linkage, one possible architecture is:
| File |
Contents |
Who May Need It? |
|
Contact file
|
Study code, name, email, telephone number, or other necessary contact details |
Authorised staff responsible for participant communication |
|
Research dataset
|
Study code and substantive research variables |
Researchers and analysts who need the study data |
|
Linkage information
|
Information necessary to connect identifiers and study records, if not already contained in the contact file |
A restricted subset of authorised personnel |
The exact architecture should match the study, applicable law, institutional systems, and security requirements. The point is functional separation: not every person who analyses responses needs access to every identifier.
Sometimes Contact Information Does Not Need to Be Linkable to Responses at All
Consider an anonymous online survey offering participants entry into a prize draw. The researcher needs email addresses to select and contact winners, but the email addresses do not need to be connected to survey answers.
One possible design is to collect survey responses first and direct interested participants to a separate contact form that does not carry a participant identifier or other mechanism connecting the contact submission to the survey response.
If the systems are genuinely separated and no other identifying information is collected, the contact list can exist without necessarily compromising the anonymity of the survey responses.
Watch Out
A second form is not automatically separate merely because it appears on another webpage. Unique URLs, embedded participant IDs, timestamps, account information, platform logs, IP addresses, or other metadata may preserve a connection. Check the actual technical data flow before promising that contact details cannot be linked to responses.
Other Studies Need the Link and Should Preserve It Deliberately
Longitudinal research may need to connect the same person's responses across multiple waves. Clinical or intervention studies may need participant-specific follow-up. Researchers may need to respond to withdrawal requests, return individual results where appropriate, verify eligibility, or connect research data with other authorised records.
In these circumstances, destroying every identity link at collection would interfere with legitimate study functions.
The solution is not to call the data anonymous. It is to retain the necessary link in a controlled manner, explain it accurately to participants, and determine who actually needs access to identifiable information.
Separation Reduces the Consequences of Some Access Events
Suppose an analyst accidentally receives the research dataset. If names and contact details are embedded in every row, the analyst immediately receives direct identifiers as well.
If those identifiers are held separately, obtaining one file may not provide the same direct identity connection. The effectiveness of this safeguard depends on what remains in the research dataset, because indirect identifiers can still make participants identifiable.
Separation therefore reduces one category of risk rather than guaranteeing non-identifiability.
Contact Information Should Have Its Own Access Rules
Researchers sometimes secure "the dataset" while leaving participant contact lists in ordinary email inboxes, spreadsheets, shared drives, calendars, or scheduling systems.
Contact information is itself personal information and should be included in the study's data-management plan. Relevant questions include:
- where it will be collected;
- where it will be stored;
- who can access it;
- how it will be transferred;
- whether it is included in backups;
- how long it is needed; and
- what happens when its purpose ends.
Separating identifiers from responses accomplishes little if the identifier file is then broadly accessible.
The Linkage Key Deserves Stronger Access Control Than Routine Analysis Data
A linkage file can turn coded research data back into participant-specific information. It should therefore not be treated as an administrative afterthought.
OHRP defines coded private information as information in which identifying information has been replaced by a code while a key exists that enables linkage. Its guidance also illustrates how access to that key can affect whether investigators can readily ascertain participants' identities under the US Common Rule.
The legal consequences vary by framework, but the operational lesson is straightforward: the ability to reconnect identity is itself sensitive research infrastructure.
Separation Can Also Support Role-Based Access
Different research personnel often need different information.
A recruitment coordinator may need participant names and contact information but not complete analytical variables. A statistician may need detailed response data but not names. A principal investigator may require authorised access to both in particular circumstances.
Separating files makes these distinctions technically possible. If everything sits in one spreadsheet, access control becomes binary: either someone gets the whole file or they do not.
This is one reason pseudonymization is recognised in data-protection guidance as an important technical and organisational safeguard even though it does not itself make personal data anonymous.
Eventually, the Identity Link May No Longer Be Necessary
Some studies need contact information during recruitment and follow-up but not indefinitely.
Once participant communication ends, the team should determine whether contact information or the linkage mechanism still serves an approved research, regulatory, integrity, archival, or other legitimate purpose. The answer depends on applicable retention requirements and the study design.
Researchers should not assume either that identifiers must be destroyed immediately or that they should be retained forever. The appropriate question is when participant identifiers can and should be permanently removed under the governing requirements.
07 · A Quick Checklist
Before Storing Contact Information With Research Responses
Check whether the identity link is actually necessary:
State why names, email addresses, telephone numbers, or other contact details are required.
Determine whether contact information needs to be linked to substantive responses at all.
If linkage is required, consider using a study code rather than embedding direct identifiers in the analysis dataset.
Store linkage information separately with access appropriate to its re-identification function.
Limit contact information and research-data access according to staff roles and legitimate study functions.
Check whether survey platforms, scheduling systems, or other services create hidden links through metadata, account information, unique URLs, or timestamps.
Describe the actual identity-linkage arrangement accurately in ethics and participant materials where required.
Plan how long contact information and linkage mechanisms need to remain available.
Verify institutional security requirements and applicable data-protection, research-ethics, and retention rules.