03 · What You Need to Know
Archiving Preserves Data, but It Also Preserves Their Ethical History
Start with provenance: where did these data actually come from?
Data provenance is the documented history of a dataset: how the information was collected, from whom, for what purpose, using which instruments and procedures, and what happened to it before it reached you.
For secondary researchers, provenance is both a methodological and an ethical issue. You may need to know whether participants consented to future reuse, whether records came from research or another context, whether identifiers were removed, whether variables were recoded, and whether particular participants or observations were excluded before archiving.
A clean downloadable file can conceal a rather untidy history. The absence of visible identifiers or consent information in the spreadsheet does not establish that those issues never existed.
Repository access does not necessarily equal ethical authorization
Being technically able to obtain a dataset does not answer whether your particular use is permitted. Archives operate under different access models. Some datasets are publicly downloadable. Others require registration, institutional affiliation, an application, a data-use agreement, ethics documentation, or approval by a data-access committee.
WHO's policy and implementation guidance on sharing and reuse of health-related data explicitly treats data reuse as requiring attention to ethical, legal, and technical issues, including privacy and appropriate data-management and sharing arrangements.
Read the repository's terms rather than inferring permission from the presence of a download button.
You can access the data
The archive's technical or administrative system has allowed you to obtain them.
You can use the data for this research
Your proposed analysis satisfies the relevant consent, ethics, legal, repository, contractual, and institutional conditions.
Original participant authorization still matters
Archived research data may have been collected under study-specific consent, broad consent, another authorized arrangement, or an older consent process that said little about future reuse. CIOMS states that institutions storing health-related research data should have governance systems for authorizing future research use and that researchers must not adversely affect the rights and welfare of the individuals from whom the data were collected.
The secondary researcher therefore needs to establish whether the proposed use is covered by the original authorization or whether another legitimate pathway applies. Whether archived data require new participant consent depends on the circumstances rather than on their archival status.
Restrictions can travel with the dataset
The original study or repository may impose restrictions concerning research topics, commercial use, attempts at re-identification, onward sharing, linkage with other datasets, geographical transfer, publication, security, or destruction after the approved project.
These are not decorative terms attached to a dataset. If you obtain access under a data-use agreement, repository license, or other formal arrangement, those conditions define what you are permitted to do. A scientifically interesting analysis does not automatically override them.
De-identification reduces some risks but does not erase every responsibility
Removing direct identifiers can substantially reduce privacy risk and may change the regulatory status of a project. It should not, however, be treated as a universal ethical exemption.
Datasets can contain unusual combinations of demographic, geographic, temporal, clinical, educational, or behavioral variables that make individuals easier to distinguish than a superficial inspection suggests. New external datasets can also create re-identification opportunities that did not exist when the archive was created.
More fundamentally, de-identification does not automatically make every secondary use ethically acceptable. Commitments to participants, repository restrictions, group harms, and inappropriate uses can remain relevant even when researchers cannot readily attach a person's name to a row.
Do not attempt re-identification simply because it is technically possible
Archived data sometimes contain enough detail to tempt a researcher to determine who particular participants are, especially when combined with publicly available information. Unless re-identification is specifically authorized and ethically justified within the research design, such attempts can violate participant expectations, repository rules, contractual obligations, or privacy requirements.
The same caution applies to linkage. Combining an archived dataset with another source can transform the informational content of both. The ethics of linking datasets without recontacting participants therefore requires its own assessment rather than being treated as ordinary secondary analysis.
Context can disappear when data are separated from their original study
Archiving often preserves variables more successfully than it preserves context. A secondary researcher may receive a codebook and dataset without having observed how questions were explained, which practical problems occurred during recruitment, how participants interpreted sensitive items, or why investigators made particular coding decisions.
This creates a risk of technically correct but substantively misleading analysis. Before interpreting archived variables, consult the original protocol, instruments, codebook, metadata, publications, data-cleaning documentation, and repository notes where available.
If the information needed to interpret a variable responsibly has been lost, uncertainty should be acknowledged rather than reconstructed from guesswork.
Secondary researchers inherit the limitations of the original design
You cannot redesign the sampling frame, add a missing confounder, change how a construct was operationalized, or repair a measurement that was unsuitable for your new purpose simply because you now possess the raw observations.
The methodological question remains whether the archived dataset can legitimately answer the different research question you want to ask. Ethical data stewardship includes resisting claims that the archived evidence cannot support.
Privacy protection continues after access is granted
Once archived data enter your possession, your responsibilities may include secure storage, access controls, approved computing environments, encryption, limits on local copies, restrictions on cloud services, destruction schedules, and breach-response procedures.
Under the Philippine Data Privacy Act and its implementing rules, personal-data processing must remain fair and lawful, compatible with declared and legitimate purposes, proportionate to those purposes, and subject to appropriate safeguards. The law also contains specific provisions concerning scientific and statistical research and confidentiality.
Repository approval therefore does not relieve the secondary research team of its own data-security responsibilities.
Data minimization still applies to secondary research
Researchers sometimes request every available variable because the dataset already exists. That logic is weak. If your study needs 20 variables, acquiring 200 may create unnecessary privacy and security exposure without improving the science.
Where applicable, request or retain only the information necessary for the approved research purpose. The Philippine implementing rules, for example, articulate proportionality by requiring processing to be adequate, relevant, suitable, necessary, and not excessive in relation to the declared purpose.
Archived data can create harms beyond individual identification
Some analyses can stigmatize communities, expose small populations, reinforce discriminatory classifications, or generate conclusions with consequences for groups even when no participant is named. These concerns can be particularly important for rare conditions, small geographic populations, Indigenous communities, marginalized groups, or datasets containing highly sensitive attributes.
Individual de-identification therefore should not exhaust the ethical analysis. Consider what can be inferred about populations as well as individuals and how findings will be communicated.
Watch Out
“Publicly available,” “archived,” “open,” and “de-identified” describe different characteristics of data. None of these labels, by itself, proves that every proposed research use is ethically or legally unrestricted.
07 · A Quick Checklist
Before Using Archived Research Data
Before beginning secondary analysis, check:
Read the repository's access conditions, license, and data-use restrictions.
Retrieve the original consent information and determine what future use participants authorized where applicable.
Review the original protocol, instruments, codebook, metadata, and relevant publications to establish data provenance.
Determine the identifiability and sensitivity of the data you will actually receive.
Request or retain only the variables reasonably necessary for your approved research purpose where practicable.
Obtain any required ethics, privacy, institutional, or data-access determination before analysis.
Implement the required security, access-control, sharing, retention, and destruction arrangements.
Check whether your planned linkage, publication, or reporting could create new identification or group-level risks.
Report the limitations inherited from the original study rather than treating the archive as data collected specifically for your question.