03 · What You Need to Know
Data sharing can extend research beyond the original participation decision
Participants often understand that researchers will collect and analyze information for the study they agreed to join. Sharing that information beyond the original research team can change the informational environment in ways that matter.
A new researcher may gain access. Data may move to another institution or repository. Future studies may ask different questions. A dataset may be retained for years. Access may be restricted to approved researchers or made publicly available.
Those differences explain why data sharing deserves explicit consideration during consent and research planning.
Start with the scope of the original consent
The first question is not simply whether the data exist. It is what participants were told could happen to them.
If the consent materials clearly explained that appropriately protected data would be deposited in a named or described repository and made available for specified future research, that agreement may support those activities, subject to the applicable requirements.
If participants were told that their identifiable information would be used only by the current research team for one particular study, later sharing for unrelated research presents a different problem.
OHRP advisory guidance concerning biospecimens and data specifically recommends examining whether a proposed secondary use falls within the scope of the original consent. When it does not, additional consent or an appropriate waiver may be necessary under the Common Rule framework.
Sharing within the research team is not necessarily the same as sharing beyond it
A participant may reasonably expect that authorized members of the research team will need access to study data.
Sharing the same data with investigators at another institution, depositing them in a repository, or making them available to future researchers can expand who has access and what the data may be used to investigate.
The consent process should therefore describe data access and sharing at a level appropriate to the research.
A generic statement such as "your information will be kept confidential" can become misleading if the protocol simultaneously plans broad distribution of participant-level data.
Identifiability changes the regulatory analysis
Whether information is identifiable is often central to the rules governing secondary research.
Under HHS regulations, identifiable private information is private information for which the participant's identity is or may readily be ascertained by the investigator or associated with the information. OHRP also recognizes circumstances in which coded information is not considered identifiable to a secondary investigator because that investigator cannot readily ascertain the participant's identity.
This distinction can affect whether a particular secondary use constitutes human-subject research under the Common Rule.
But regulatory identifiability and practical privacy risk should not be treated as identical concepts.
Removing names does not automatically make data anonymous
A dataset can contain no names and still permit re-identification.
Dates, locations, rare characteristics, occupations, combinations of demographic variables, genomic information, free-text responses, or linked datasets may allow individuals to be recognized directly or indirectly.
Researchers should therefore avoid using "anonymous," "de-identified," and "coded" interchangeably.
Identifiable data
Information from which a participant's identity is or may readily be ascertained under the applicable framework.
Coded data
Direct identifiers are replaced by a code, but a key or other mechanism may still connect the information to the participant.
De-identified or nonidentifiable data
Information processed so that the relevant investigator cannot readily identify the individual under the applicable standard.
The precise definitions vary across laws and research frameworks, so researchers should use the terminology applicable to their data rather than relying on informal labels.
De-identification can change regulatory requirements without erasing prior commitments
This is a subtle but important distinction.
OHRP guidance explains that some secondary research involving coded private information or biospecimens may not constitute human-subject research when the secondary investigators cannot readily ascertain participants' identities. In those circumstances, Common Rule consent requirements may not apply to that secondary research.
Yet OHRP advisory guidance also emphasizes that researchers and institutions should honor commitments made to participants in the original consent. For example, if participants were told that their specimens would be used only for research on one condition, later de-identification does not ethically make that promise disappear merely because the secondary activity falls outside particular human-subject regulations.
Regulatory permission and fidelity to participant agreements are therefore related but not identical.
Broad consent can authorize some future secondary research
The revised U.S. Common Rule created a specific regulatory option called broad consent for storage, maintenance, and secondary research use of identifiable private information and identifiable biospecimens.
Broad consent is not simply a sentence saying, "Your data may be used for future research."
It contains specified elements. Among other things, participants receive a general description of the types of research that may be conducted, a description of the information or biospecimens involved, whether sharing may occur and the types of institutions or researchers that may conduct future research, and relevant information about how long materials may be stored and used.
Broad consent is one regulatory option rather than a universal requirement for all future research. Other pathways for secondary research may also be available depending on the circumstances.
Broad consent does not mean unlimited consent
The word "broad" can invite overinterpretation.
Under the Common Rule, future secondary research relying on broad consent still needs to fall within the scope of what participants were told. For the exemption at 45 CFR 46.104(d)(8), limited IRB review includes determining that the proposed secondary research is within the scope of the broad consent obtained.
OHRP advisory guidance also notes that a bare checkbox permitting unspecified future research may be insufficient to provide meaningful information about future use.
Broad consent is therefore broader than study-specific consent, not boundaryless permission.
A participant's refusal of broad consent can matter later
The revised Common Rule contains a particularly consequential rule: when an individual was asked to provide broad consent for storage, maintenance, and secondary research use of identifiable private information or identifiable biospecimens and refused, an IRB cannot later waive consent for the storage, maintenance, or secondary research use covered by that refusal.
This prevents an explicit refusal from simply being bypassed through the ordinary waiver mechanism.
Researchers considering broad consent should therefore understand both what agreement permits and what refusal can mean for future research options.
Public repositories and controlled-access repositories create different exposures
Not all data sharing is equally open.
Open or public access
Data are made broadly accessible with few or no case-by-case restrictions on who can obtain them.
Controlled access
Researchers or institutions must satisfy specified conditions before gaining access, which may include an application, data-use agreement, research-purpose restrictions, security requirements, or oversight.
Controlled access can reduce some privacy and misuse risks, but it does not mean the data have not been shared. Participants may still reasonably want to know that researchers outside the original team could gain access.
The appropriate access model depends on the sensitivity and identifiability of the data, consent obtained, repository requirements, applicable law, and scientific goals.
Data sharing plans should be considered before consent is written
A common mistake is to develop a data-sharing plan only when a journal, funder, or repository asks for one near the end of the project.
By then, the consent form may promise that only the research team will see the data.
The better approach is to decide during study design whether participant-level data may be retained, deposited, shared, linked, or reused and then make the consent information consistent with those plans.
This is especially important when research funders or journals have data-sharing expectations.
A funder or journal data-sharing policy does not automatically override participant consent
Open-science and data-sharing policies can promote reproducibility and further research. They do not create unlimited authority to disregard promises made to participants or applicable privacy and research requirements.
If participant consent or legal restrictions prohibit a particular form of sharing, researchers need to determine what sharing is actually permissible. That may mean using a controlled-access mechanism, sharing a more strongly de-identified dataset, providing metadata rather than participant-level data, or explaining legitimate restrictions when a policy permits exceptions.
Planning for these issues prospectively is considerably easier than discovering them after publication acceptance.
Secondary research may sometimes proceed without new consent
Not every future use requires researchers to recontact every participant.
Depending on the data and governing framework, secondary research may fall outside the definition of human-subject research, qualify for an exemption, fall within valid broad consent, or proceed under an IRB-approved waiver of consent.
For example, OHRP guidance explains that certain secondary research involving coded information may not involve human subjects under the Common Rule when investigators cannot readily ascertain the individuals' identities.
For identifiable information or biospecimens, other pathways may apply, including particular exemptions and waivers under the revised Common Rule.
The correct conclusion is therefore not "new consent is always required." It is "the original participation consent does not by itself answer every future sharing question."
Data sharing can be more consequential for some data types
Genomic data, detailed geolocation, rare-disease information, linked administrative records, qualitative transcripts, photographs, video, and highly detailed longitudinal datasets can create distinctive privacy or re-identification concerns.
A dataset that appears harmless in isolation may become more revealing when combined with other available information.
Researchers should therefore assess the actual disclosure risk rather than applying the same sharing strategy to every dataset merely because direct names have been removed.
Qualitative data deserve special care
Sharing an interview transcript is not equivalent to sharing a spreadsheet of numerical measurements.
Transcripts may contain names, workplaces, relationships, distinctive events, locations, personal histories, or combinations of details that make participants recognizable. Removing direct identifiers can also alter the meaning of qualitative material.
Researchers planning to archive qualitative data should consider these issues during consent and data-management design, including whether full transcripts, redacted transcripts, excerpts, or controlled access are appropriate.
Publication and data sharing are different acts
A journal article can report aggregated results without giving anyone access to the underlying participant-level dataset.
Data sharing, by contrast, may allow other researchers to inspect or analyze the underlying information.
That is why consent concerning publication of research results and consent concerning data sharing should be analyzed separately.
International sharing can add another layer
Sending participant data to researchers in another country may trigger additional legal, contractual, institutional, or data-protection requirements.
The relevant rules depend on where data originate, where they are transferred, their identifiability and sensitivity, and which laws or agreements apply.
Researchers should therefore avoid promising international data sharing before confirming that the proposed transfers can actually be conducted lawfully and consistently with participant consent.
Withdrawal from sharing has practical limits
A participant may later ask that their data no longer be shared.
Researchers may be able to stop future distribution from a repository or research database, depending on the consent terms and system. But data already distributed to other researchers, incorporated into completed analyses, or included in archived datasets may not always be recoverable.
The consent information should explain relevant limits rather than promising that data can always be retrieved from every recipient.
Watch Out
Do not promise participants that their data will "never leave the research team" if your protocol, repository plan, collaboration agreements, funder requirements, or future-research plans anticipate sharing. Consent language and actual data governance should describe the same research.