01 · The Question
Has Someone Actually Tested the Finding Again With New Data?
A finding can be cited hundreds of times without ever being independently replicated. Conversely, a later study may test essentially the same proposition without using the word “replication” anywhere in its title or abstract.
This makes replication surprisingly difficult to identify during literature synthesis. Searching for papers labeled “replication study” is not enough. You need to determine whether subsequent research collected new data capable of providing a meaningful new test of the earlier conclusion.
The important question is therefore not simply whether later researchers obtained a similar result. It is whether independent evidence tested the same scientific claim and produced results sufficiently consistent with it, given the uncertainty and relevant differences between studies.
03 · What You Need to Know
Replication Tests Whether a Finding Survives New Data
Terminology around reproducibility and replication varies among disciplines. The U.S. National Academies explicitly acknowledges this inconsistency and adopts a useful distinction: reproducibility means obtaining consistent computational results using the same input data and computational procedures, while replicability means obtaining consistent results across studies addressing the same scientific question when each obtains its own data.
That distinction is useful for literature synthesis because it separates two important questions. Can the original result be regenerated from the original evidence? And does the scientific finding appear again when new evidence is collected?
Reproducibility
Can consistent computational results be obtained from the same data and analytical procedures?
Replicability
Do studies using newly obtained data produce results consistent with the same scientific claim?
Independent Replication Requires New Evidence
If another team downloads the original dataset and reruns the analysis, that can provide an important reproducibility check. It may uncover coding errors, undocumented analytical choices, or discrepancies in reported results.
It is not, under the National Academies' terminology, an empirical replication because the same observations are being reused. A replication collects new data relevant to the same scientific question.
For evidence synthesis, record this distinction explicitly. Computational reproducibility and empirical replication strengthen confidence in different ways.
Independence Is a Matter of Degree
A replication performed by the original research team using a new sample provides new data. A replication by a different team in another institution adds investigator independence. A study using a different population or operationalization may additionally test generalizability or robustness.
These are not interchangeable achievements. When describing a replicated conclusion, specify what changed and what remained the same.
| Later study |
What it can test |
Important limitation |
| Reanalysis of original data |
Computational reproducibility or analytical robustness |
No new empirical observations |
| New sample, original research team |
Whether the result recurs in newly collected data |
Some investigator or procedural dependencies may remain |
| New sample, independent team, closely similar procedure |
Independent replication under similar conditions |
Generalizability beyond those conditions may remain unknown |
| Independent team with meaningful methodological changes |
Whether the claim survives a broader empirical test |
Differences may complicate attribution when results diverge |
| New population or setting |
Replication plus evidence relevant to generalizability |
Population differences may legitimately alter the effect |
A Replication Does Not Need an Identical Numerical Result
Independent studies will rarely produce identical estimates. Sampling variability alone guarantees some difference, and real changes in populations, settings, implementation, or measurement may add further variation.
The National Academies emphasizes that replication must be assessed with uncertainty in mind. Consistency may involve the proximity of results together with the precision of their estimates rather than exact numerical equality.
Consequently, deciding whether two studies replicate should not be reduced to asking whether both produced a P value below.05. Two studies can produce similar effect estimates while falling on opposite sides of a significance threshold. Conversely, two statistically significant results can differ substantially in magnitude.
Replication Is About the Claim, Not Merely the Procedure
A close or direct replication attempts to test a previous result under conditions where there is no prior reason to expect a meaningfully different outcome. Yet no replication can reproduce every feature of an earlier study perfectly. Participants, time, location, researchers, and numerous small procedural details inevitably differ.
This is why replication ultimately concerns what those similarities and differences mean for the claim being tested. Nosek and Errington propose a particularly useful claim-centered formulation: a replication should be capable of changing confidence in the prior claim regardless of which way the new result turns out.
If a positive result would be celebrated as confirmation but a negative result would automatically be dismissed as “not a real replication,” the claim has not been exposed to much evidential risk.
One Successful Replication Increases Confidence but Does Not Settle the Matter
The National Academies states explicitly that a successful replication does not guarantee that an original scientific result was correct. Both studies could share an unnoticed bias, use the same problematic measurement, or happen to produce compatible estimates despite an incorrect underlying interpretation.
Still, a credible replication matters. The original finding is no longer an isolated observation. Confidence may increase further when replication occurs across genuinely independent teams, datasets, methods, and settings.
At that point, the conclusion may begin to be supported by several independent lines of evidence rather than replication alone.
One Failed Replication Does Not Automatically Refute the Original Finding
The reverse mistake is equally problematic. A replication may differ from the original study in consequential ways. Either study may be imprecise. The phenomenon may be context-sensitive. Implementation may vary. Random sampling variation can also produce different estimates.
The National Academies therefore cautions that a single failed replication does not conclusively refute an original claim.
Instead, inspect the original and replication studies symmetrically. Ask about their designs, sample sizes, uncertainty, measurements, implementation, populations, and risks of bias. The original study does not receive automatic methodological privilege merely because it came first.
Repeated Replication Failures Are More Informative Than One
If several credible independent attempts repeatedly fail to obtain results compatible with an original claim, confidence in that claim should generally decrease. The National Academies notes that repeated findings of comparable results tend to strengthen confidence, while repeated failures to confirm an earlier result cast doubt on the original conclusion.
At the same time, persistent differences can be scientifically productive. They may reveal a previously unnoticed moderator, population boundary, implementation requirement, or measurement problem. A replication failure is therefore not merely a red mark beside an original paper. Sometimes it tells you that the original claim was too broad.
Replication and Generalizability Answer Different Questions
A finding replicated repeatedly under highly similar conditions may be reliable within those conditions without being broadly generalizable.
Conversely, studies conducted in substantially different populations may produce somewhat different estimates while revealing a coherent conditional pattern. The National Academies distinguishes replicability from generalizability, with the latter concerning whether results apply to contexts or populations different from those originally studied.
To claim that a conclusion is robust across populations and methods, you need broader evidence than repeated close replication alone.
Watch Out
Do not label a finding “replicated” merely because two studies produced statistically significant results in the same direction. Compare the estimates, their uncertainty, the scientific questions addressed, and meaningful differences between the studies.
Replications May Not Announce Themselves
The National Academies notes that many studies effectively replicate previous work without being explicitly labeled as replications. Researchers may repeat an important result as part of a larger investigation, or a field may independently conduct closely related studies without describing them as formal replication attempts.
For a literature review, this means searching by the substantive claim as well as by terms such as “replication.” Examine later studies to determine whether their data provide a genuine new test of the proposition.
06 · What This Means for You
Evaluate Replication at the Level of the Claim
Start with a specific conclusion rather than a paper. Then identify studies that collected new data capable of testing that conclusion. For each one, record who conducted it, what changed, what remained similar, and how its estimate compares with the earlier evidence.
A simple decision framework
If a later paper only reruns the original data
Treat it as evidence about reproducibility or analytical robustness, not an independent new-data replication.
If a new study collects independent data to test the same claim
Compare its result and uncertainty with the original rather than comparing significance labels.
If several credible independent studies produce compatible results
Increase confidence in the conclusion while preserving any uncertainty about effect magnitude and scope.
If one replication attempt disagrees
Investigate uncertainty, design, population, implementation, and measurement differences before treating either study as decisive.
If credible independent studies repeatedly disagree with the original finding
Reduce confidence in the original conclusion and investigate whether a narrower or context-dependent explanation better fits the evidence.
When reporting replication in your synthesis, be concrete. “The finding has been replicated” is less informative than stating that three independent teams obtained compatible effects in newly collected samples using closely similar procedures.
That wording tells readers what kind of replication evidence actually exists rather than turning replication into a ceremonial badge of scientific respectability.
07 · A Quick Checklist
Check Whether a Conclusion Has Really Been Replicated
For each claimed replication, check:
Did the later study collect new data rather than only reanalyze the original dataset?
Does it genuinely test the same scientific claim?
Was the study conducted by an independent research team, and if not, what dependencies remain?
Are the effect estimates compatible when their uncertainty is considered?
Am I comparing results rather than simply comparing whether P values crossed a significance threshold?
Were there consequential differences in population, setting, measurement, procedure, or implementation?
If results diverge, have I evaluated both the original and replication studies rather than assuming the original is the benchmark?
Are multiple replication attempts independent of one another?
Have I kept replication separate from the broader question of generalizability?