Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Ethical Responsibilities Come With Using Archived Research Data?

Archived research data may be scientifically valuable long after the original study ends, but depositing data in an archive does not erase researchers' ethical responsibilities. Secondary users need to understand the data's provenance, authorization, restrictions, privacy risks, and governance before reuse.

368
Ethics of Using Archived Research Data Guide 368 of 398
01 · The Question

Does Archiving Research Data Make Them Free to Use?

A dataset may sit in a university repository, disciplinary archive, research center, or long-term project database years after the original study has ended. You find it, obtain access, and realize that it could answer your research question without recruiting a single new participant.

It is tempting to think that the difficult ethical work happened during the original study. But archiving changes where the data are stored, not necessarily what researchers owe the people represented in them. Ethical responsibilities can follow data into their second, third, or twentieth research life.

02 · The Short Answer

Archived Data Still Carry Ethical Conditions and Responsibilities

In Brief

Using archived research data can be ethically appropriate, but researchers remain responsible for respecting the original authorization, access conditions, privacy and confidentiality, data provenance, applicable ethics review, and the limits of what the archived data can support.

An archive is not an ethical reset button. Before reuse, determine where the data came from, what participants were told, how the data were transformed, what restrictions remain attached, and whether your proposed study is permitted under the applicable ethical, legal, institutional, and repository framework.

03 · What You Need to Know

Archiving Preserves Data, but It Also Preserves Their Ethical History

Start with provenance: where did these data actually come from?

Data provenance is the documented history of a dataset: how the information was collected, from whom, for what purpose, using which instruments and procedures, and what happened to it before it reached you.

For secondary researchers, provenance is both a methodological and an ethical issue. You may need to know whether participants consented to future reuse, whether records came from research or another context, whether identifiers were removed, whether variables were recoded, and whether particular participants or observations were excluded before archiving.

A clean downloadable file can conceal a rather untidy history. The absence of visible identifiers or consent information in the spreadsheet does not establish that those issues never existed.

Repository access does not necessarily equal ethical authorization

Being technically able to obtain a dataset does not answer whether your particular use is permitted. Archives operate under different access models. Some datasets are publicly downloadable. Others require registration, institutional affiliation, an application, a data-use agreement, ethics documentation, or approval by a data-access committee.

WHO's policy and implementation guidance on sharing and reuse of health-related data explicitly treats data reuse as requiring attention to ethical, legal, and technical issues, including privacy and appropriate data-management and sharing arrangements.

Read the repository's terms rather than inferring permission from the presence of a download button.

You can access the data The archive's technical or administrative system has allowed you to obtain them.
You can use the data for this research Your proposed analysis satisfies the relevant consent, ethics, legal, repository, contractual, and institutional conditions.

Original participant authorization still matters

Archived research data may have been collected under study-specific consent, broad consent, another authorized arrangement, or an older consent process that said little about future reuse. CIOMS states that institutions storing health-related research data should have governance systems for authorizing future research use and that researchers must not adversely affect the rights and welfare of the individuals from whom the data were collected.

The secondary researcher therefore needs to establish whether the proposed use is covered by the original authorization or whether another legitimate pathway applies. Whether archived data require new participant consent depends on the circumstances rather than on their archival status.

Restrictions can travel with the dataset

The original study or repository may impose restrictions concerning research topics, commercial use, attempts at re-identification, onward sharing, linkage with other datasets, geographical transfer, publication, security, or destruction after the approved project.

These are not decorative terms attached to a dataset. If you obtain access under a data-use agreement, repository license, or other formal arrangement, those conditions define what you are permitted to do. A scientifically interesting analysis does not automatically override them.

De-identification reduces some risks but does not erase every responsibility

Removing direct identifiers can substantially reduce privacy risk and may change the regulatory status of a project. It should not, however, be treated as a universal ethical exemption.

Datasets can contain unusual combinations of demographic, geographic, temporal, clinical, educational, or behavioral variables that make individuals easier to distinguish than a superficial inspection suggests. New external datasets can also create re-identification opportunities that did not exist when the archive was created.

More fundamentally, de-identification does not automatically make every secondary use ethically acceptable. Commitments to participants, repository restrictions, group harms, and inappropriate uses can remain relevant even when researchers cannot readily attach a person's name to a row.

Do not attempt re-identification simply because it is technically possible

Archived data sometimes contain enough detail to tempt a researcher to determine who particular participants are, especially when combined with publicly available information. Unless re-identification is specifically authorized and ethically justified within the research design, such attempts can violate participant expectations, repository rules, contractual obligations, or privacy requirements.

The same caution applies to linkage. Combining an archived dataset with another source can transform the informational content of both. The ethics of linking datasets without recontacting participants therefore requires its own assessment rather than being treated as ordinary secondary analysis.

Context can disappear when data are separated from their original study

Archiving often preserves variables more successfully than it preserves context. A secondary researcher may receive a codebook and dataset without having observed how questions were explained, which practical problems occurred during recruitment, how participants interpreted sensitive items, or why investigators made particular coding decisions.

This creates a risk of technically correct but substantively misleading analysis. Before interpreting archived variables, consult the original protocol, instruments, codebook, metadata, publications, data-cleaning documentation, and repository notes where available.

If the information needed to interpret a variable responsibly has been lost, uncertainty should be acknowledged rather than reconstructed from guesswork.

Secondary researchers inherit the limitations of the original design

You cannot redesign the sampling frame, add a missing confounder, change how a construct was operationalized, or repair a measurement that was unsuitable for your new purpose simply because you now possess the raw observations.

The methodological question remains whether the archived dataset can legitimately answer the different research question you want to ask. Ethical data stewardship includes resisting claims that the archived evidence cannot support.

Privacy protection continues after access is granted

Once archived data enter your possession, your responsibilities may include secure storage, access controls, approved computing environments, encryption, limits on local copies, restrictions on cloud services, destruction schedules, and breach-response procedures.

Under the Philippine Data Privacy Act and its implementing rules, personal-data processing must remain fair and lawful, compatible with declared and legitimate purposes, proportionate to those purposes, and subject to appropriate safeguards. The law also contains specific provisions concerning scientific and statistical research and confidentiality.

Repository approval therefore does not relieve the secondary research team of its own data-security responsibilities.

Data minimization still applies to secondary research

Researchers sometimes request every available variable because the dataset already exists. That logic is weak. If your study needs 20 variables, acquiring 200 may create unnecessary privacy and security exposure without improving the science.

Where applicable, request or retain only the information necessary for the approved research purpose. The Philippine implementing rules, for example, articulate proportionality by requiring processing to be adequate, relevant, suitable, necessary, and not excessive in relation to the declared purpose.

Archived data can create harms beyond individual identification

Some analyses can stigmatize communities, expose small populations, reinforce discriminatory classifications, or generate conclusions with consequences for groups even when no participant is named. These concerns can be particularly important for rare conditions, small geographic populations, Indigenous communities, marginalized groups, or datasets containing highly sensitive attributes.

Individual de-identification therefore should not exhaust the ethical analysis. Consider what can be inferred about populations as well as individuals and how findings will be communicated.

Watch Out

“Publicly available,” “archived,” “open,” and “de-identified” describe different characteristics of data. None of these labels, by itself, proves that every proposed research use is ethically or legally unrestricted.

04 · A Practical Example

When a Publicly Discoverable Dataset Still Requires Ethical Work

Hypothetical Example

Reusing a decade-old longitudinal dataset

A researcher discovers an archived dataset from a longitudinal study of university students conducted ten years earlier. The archive contains demographic information, survey responses, academic measures, a detailed codebook, and documentation describing restricted access for approved research.

Check provenance The researcher reads the original protocol, questionnaires, consent materials, metadata, and publications to understand how the variables were created and what participants were told.
Check authorization The proposed question is compared with the original consent and repository conditions rather than assuming that archival deposit created unrestricted permission.
Request only what is needed The researcher does not request every archived variable merely because they are available. Variables unnecessary for the approved question are excluded where practicable.
Protect the working copy The dataset is stored and analyzed according to the repository's security and access requirements, and it is not redistributed to colleagues who were not authorized to receive it.
Interpret with context The researcher reports limitations created by the original sampling and measurement design and avoids treating archived variables as though they had been collected specifically for the new study.
05 · What Researchers Often Get Wrong

Common Ethical Mistakes When Working With Archived Data

Misconception

“If a repository gave me the dataset, the ethical review is finished.”

Repository access and research-ethics requirements are not necessarily the same thing. Your institution or jurisdiction may require an additional determination concerning human-subjects research, consent, privacy, or data protection.

Misconception

“Old data have fewer ethical obligations.”

Age does not automatically extinguish confidentiality, consent restrictions, contractual obligations, or privacy risks. In some cases, technological advances can make old data more revealing than they were when collected.

Misconception

“If the dataset has no names, privacy is no longer relevant.”

Direct identifiers are only one route to identification. Detailed combinations of attributes and linkage with external sources may create residual risk, and ethical concerns can remain even when individual identification is unlikely.

Misconception

“Open data means I can do anything with them.”

Open access may remove some access barriers, but researchers still need to examine licensing, consent, ethical context, privacy implications, and the conditions under which the archive released the data.

Misconception

“I do not need the original study documentation because I only need the numbers.”

Without provenance and metadata, you may misunderstand variables, coding, missingness, eligibility, timing, or sampling. Responsible secondary analysis requires understanding how the numbers became numbers in the first place.

06 · What This Means for You

Treat Archived Data as Entrusted Research Material, Not Found Property

Before opening your statistical software, reconstruct the dataset's ethical and methodological history. The key question is not merely whether you can obtain the file, but whether you understand enough about its origin and conditions to use it responsibly.

A simple decision framework

If provenance, consent, and repository conditions are well documented
Compare your proposed study with those conditions and complete the applicable ethics, privacy, and access processes.
If access is restricted
Use the repository's authorized application process and comply with its data-use and security conditions.
If crucial consent or provenance information is missing
Seek clarification from the archive, original custodians, or appropriate institutional authority rather than filling the gap with assumptions.
If the proposed analysis requires linkage, re-identification, or a materially different use
Treat that as an additional ethical and governance question rather than assuming the original archive approval covers it.
07 · A Quick Checklist

Before Using Archived Research Data

Before beginning secondary analysis, check:
Read the repository's access conditions, license, and data-use restrictions.
Retrieve the original consent information and determine what future use participants authorized where applicable.
Review the original protocol, instruments, codebook, metadata, and relevant publications to establish data provenance.
Determine the identifiability and sensitivity of the data you will actually receive.
Request or retain only the variables reasonably necessary for your approved research purpose where practicable.
Obtain any required ethics, privacy, institutional, or data-access determination before analysis.
Implement the required security, access-control, sharing, retention, and destruction arrangements.
Check whether your planned linkage, publication, or reporting could create new identification or group-level risks.
Report the limitations inherited from the original study rather than treating the archive as data collected specifically for your question.
08 · Frequently Asked Questions

Frequently Asked Questions About Archived Research Data

Do archived data always count as secondary data?

When researchers use existing archived information for a new analysis rather than collecting it directly for that study, the activity is generally described as secondary data use or secondary analysis. The precise regulatory classification depends on the applicable framework.

Do I need ethics approval to analyze archived data?

Possibly. Requirements depend on factors such as identifiability, the source and nature of the information, the proposed use, jurisdiction, and institutional policy. OHRP recommends that institutions designate an authorized person or entity to determine whether secondary research involving coded information constitutes exempt or nonexempt human-subjects research.

Can I download an open dataset and share it with my research team?

Check the repository's terms, license, consent restrictions, and privacy conditions first. Public access does not necessarily mean unrestricted redistribution.

Can I combine archived data with another dataset?

Potentially, but linkage may create new information and privacy risks. Verify whether linkage is permitted by the original authorization, repository conditions, applicable law, and ethics or institutional review requirements.

What if I cannot find the original consent form?

Do not assume that missing documentation means unrestricted use. Contact the repository or data custodian and seek the appropriate institutional determination concerning what can legitimately be established about the original authorization.

Can archived data be retained forever by the secondary researcher?

Not automatically. Retention may be governed by the repository agreement, consent, ethics approval, privacy law, institutional policy, or other requirements. Long-term archival retention by an authorized repository and indefinite possession of a working copy by every secondary researcher are not the same arrangement.

Should I cite the dataset itself?

Where the repository provides a persistent identifier or required citation, follow its citation instructions in addition to citing relevant publications. Proper data citation supports provenance, credit, and reproducibility.

09 · The Bottom Line

Archived Data Remain Connected to the People and Conditions That Produced Them

The Bottom Line

Using archived research data carries continuing responsibilities to understand their provenance, respect the authorization and restrictions attached to them, protect privacy and confidentiality, and use the data only in ways that are methodologically and ethically defensible.

Treat the archive as a custodian of research material, not as a mechanism that wipes away its history. Responsible reuse begins by finding out what happened before the dataset reached you and what conditions still govern what happens next.

10 · Sources and Further Reading

Authoritative Guidance on Archived Data and Secondary Research

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes