01 · The Question
Can You Replicate a Study If You Cannot Access the Original Dataset?
You have identified a study worth replicating, but the original dataset is nowhere to be found. There is no public repository link, no supplementary data file, and perhaps no indication that the data were ever made available. You may have contacted the authors and learned that the dataset can no longer be shared or retrieved.
Does that prevent replication?
Usually, no. A replication normally involves collecting new data to test a claim from previous research. The original dataset can be highly valuable for verifying analyses, planning the new study, understanding unexpected details, and comparing results, but access to it is not generally a prerequisite for conducting a replication.
The complication is that unavailable data sometimes signal a broader transparency problem. If essential procedures, materials, variable definitions, analytical decisions, or other information are also unavailable, you may no longer know enough about the original study to design an interpretable replication.
03 · What You Need to Know
Missing Original Data Affect Reproducibility and Replication Differently
The distinction between reproducibility and replicability is particularly useful here. Terminology varies across disciplines, but the National Academies defines reproducibility as obtaining consistent computational results using the same input data, computational steps, methods, code, and conditions of analysis. Replicability, by contrast, concerns obtaining consistent results across studies aimed at answering the same scientific question, each using its own data.
Under those definitions, unavailable original data can make computational reproduction impossible while leaving replication possible.
Reproducing the analysis
You need the original input data, together with sufficient information about the computational and analytical procedures, to determine whether the reported result can be recomputed.
Replicating the finding
You collect new data and conduct a new study capable of providing evidence about the claim from the original research.
This distinction prevents a common mistake: assuming that inability to rerun the original analysis means the scientific claim cannot be tested again.
What Do You Actually Need to Conduct a Replication?
You need enough information to identify what was tested and to reconstruct the features of the study that matter for interpreting the claim. Depending on the research design, this may include participant eligibility criteria, sampling procedures, intervention or exposure details, measures, materials, timing, experimental conditions, variable definitions, exclusion rules, preprocessing procedures, and statistical analyses.
Not every replication requires every original detail to be identical. Nosek and Errington emphasize that exact replication is impossible because original and replication studies inevitably differ in some respects. The challenge is determining which conditions are presumed necessary for an informative test of the original claim.
This means the absence of raw data by itself may be manageable. The absence of raw data combined with an incomplete methods section, unavailable instruments, undocumented analytical decisions, and inaccessible materials can be much more serious.
Why Are the Original Data Still Valuable?
Although replication generates new data, access to the original dataset can improve the process considerably. You may be able to verify reported descriptive statistics, inspect variable coding, understand exclusions, reproduce analytical decisions, examine distributions, estimate plausible effect sizes, or identify discrepancies between the written methods and the implemented analysis.
The data can also help distinguish two questions that otherwise become entangled: whether the original reported result follows from the original data and analysis, and whether new data provide evidence consistent with the underlying scientific claim.
Without the dataset, the first question may remain unresolved even if you successfully complete the second.
Unavailable Data Do Not Automatically Mean Poor Research Practice
Do not assume misconduct or inadequate research practice simply because data are unavailable. There are legitimate reasons why original data cannot be shared publicly or indefinitely.
Human-participant data may be restricted by consent agreements, privacy protections, legal obligations, data-use agreements, institutional policies, or risks of re-identification. Proprietary datasets may be subject to licensing restrictions. Historical studies may predate contemporary expectations for repository-based sharing. Files can also become inaccessible because of obsolete systems, personnel changes, or retention limits.
The relevant question for your replication is what information can lawfully and practically be obtained, not why the absence of a downloadable dataset should automatically be interpreted as suspicious.
Watch Out
“Data not publicly available” and “data do not exist” are not equivalent. Before concluding that the original data are unavailable, check the article, supplementary materials, repository records, data-availability statement, and any controlled-access procedures identified by the authors or institution.
Contacting the Original Authors Can Resolve More Than the Data Problem
If important information is missing, contacting the original authors can be useful even when they cannot provide the dataset itself. They may be able to clarify procedures, supply questionnaires or stimuli, explain coding decisions, identify software settings, describe implementation details, or point you to an archive that was not linked prominently in the publication.
Replication initiatives have demonstrated how consequential such information can be. The Reproducibility Project: Cancer Biology contacted original authors for methodological details and sought original reagents, protocols, and data when developing replication designs. Later reports from the project documented barriers involving incomplete methodological documentation and failures to obtain original data, reagents, and other materials.
Author contact should nevertheless supplement, not silently replace, the published record. Document consequential clarifications so readers understand how your replication protocol was constructed.
Sometimes the Missing Data Prevent a Precise Statistical Comparison
You may be able to replicate the study but still lack information required for particular comparisons. For example, a paper may report only a significance test without sufficient descriptive statistics or uncertainty estimates. Without the original data, reconstructing the original effect estimate may be difficult or impossible.
That does not necessarily prevent a new test of the scientific claim, but it may constrain meta-analysis, equivalence testing, effect-size comparison, or other analyses requiring information that was not reported.
Be explicit about this limitation. Do not manufacture missing statistics from incomplete published information.
What If the Methods Are Also Incomplete?
This is where the problem becomes more consequential. If you cannot determine how participants were selected, how an intervention was delivered, how an outcome was calculated, which observations were excluded, or which analysis produced the published result, you may not know what conditions you are attempting to reproduce.
You can still conduct a new study addressing the broader scientific question, but its status as a close replication becomes harder to defend. The issue then resembles the broader problem of being unable to reproduce the original study closely.
Do Not Reconstruct Unknown Methods by Guessing
If an essential detail cannot be recovered, distinguish clearly among what the paper reports, what the authors later clarified, and what you decided for the replication.
For example, if the publication does not specify how missing observations were handled, do not quietly select a method and imply that it reproduces the original analysis. State that the original procedure could not be determined, explain the approach you adopted, and discuss how that uncertainty affects comparison.
Transparent deviation is methodologically preferable to invented fidelity.
Missing Original Data Can Itself Reveal a Limitation in the Evidence Base
When a consequential finding cannot be computationally checked because the original data and analytical record are unavailable, independent replication may become more valuable, not less. New evidence cannot reconstruct the lost dataset, but it can reduce the field's dependence on an empirical claim that cannot be independently recomputed.
This can be particularly important when the study has never been independently replicated. The justification should remain carefully framed: the problem is limited verifiability and independent evidence, not an assumption that unavailable data make the published finding false.
07 · A Quick Checklist
Before Replicating a Study With Unavailable Original Data
Before deciding that missing data prevent replication, check:
Read the study's data-availability statement and supplementary materials carefully.
Search repositories or archives explicitly identified by the publication, authors, journal, or project.
Determine whether the data are genuinely unavailable or merely subject to controlled-access requirements.
Separate information needed to reproduce the original analysis from information needed to conduct a new replication.
Inventory the methods, measures, materials, protocols, variable definitions, exclusion rules, and analytical details needed for your study.
Contact the original authors when consequential methodological information cannot be recovered from the published record.
Document which original details remain unknown and which methodological decisions you made independently.
State explicitly how missing original data limit comparison and interpretation of the replication.