01 · The Question
What Should You Do After an Original Finding Fails to Replicate?
The original study reports a clear finding. A later research team attempts to replicate it and obtains a substantially different result. Now the literature contains a disagreement rather than a straightforward claim.
Should you conduct another replication?
Possibly, but the rationale has changed. Before the failed replication, the question may have been whether independent evidence supports the original finding. Afterward, the more informative question is often why the studies disagree and what additional evidence could distinguish among the plausible explanations.
A third study that simply repeats one side without addressing the source of disagreement can leave the field with three results instead of two and not much more understanding.
03 · What You Need to Know
A Failed Replication Changes the Question Rather Than Ending It
The National Academies of Sciences, Engineering, and Medicine explicitly cautions that a single failed replication does not conclusively refute an original scientific claim. Non-replicability can arise for many reasons, including inherent variability, uncontrolled complexity, methodological problems, chance, or previously unrecognized phenomena.
At the same time, a failed replication should not simply be ignored. It is new evidence. The task is to determine how much that evidence should change confidence in the claim and what investigation, if any, is needed next.
“Failed to Replicate” Is Not a Single Scientific Outcome
The phrase can conceal substantial differences in what actually happened. One replication may estimate an effect near the original estimate but with wide uncertainty. Another may produce an estimate close to zero with high precision. A third may find an effect in the opposite direction. Those results do not provide identical evidence.
The National Academies notes that replicability assessment need not produce a binary pass-or-fail judgment. The degree to which results are consistent can depend on the criteria used, the uncertainty surrounding the estimates, and the scientific question being evaluated.
Before planning another study, inspect the estimates and uncertainty rather than relying only on labels such as “successful” and “failed.”
Compare the Original Study and the Failed Replication Symmetrically
Researchers sometimes assume that the original study establishes the benchmark and that the replication must explain any deviation from it. That is not necessarily justified.
Both studies are empirical observations subject to sampling variation, methodological limitations, implementation problems, and analytical choices. The original study may have overestimated an effect. The replication may have implemented an essential procedure poorly. Both could be reasonable studies conducted under conditions where the phenomenon genuinely differs.
A useful investigation therefore asks what each study contributes rather than automatically privileging whichever result appeared first.
Watch Out
Do not diagnose a failed replication solely by searching for flaws in the replication while treating the original study as error-free. Apply comparable methodological scrutiny to both.
Check Whether the Replication Tested the Same Claim Closely Enough
No two studies are literally identical. Differences in population, setting, materials, timing, measurement, implementation, and analysis are inevitable. Some differences are irrelevant to the claim; others may be consequential.
When a replication produces a discrepant result, examine whether departures from the original design could plausibly matter. This does not mean inventing a post hoc explanation for every difference. A proposed moderator or boundary condition should be scientifically defensible and, ideally, testable.
If substantial changes were introduced, you may need to revisit whether the earlier attempt was a sufficiently close replication or a more conceptual test. Understanding what kind of replication question was actually asked becomes particularly important when interpreting disagreement.
Ask Whether Sampling Variation Could Plausibly Explain the Difference
Different samples produce different estimates even when studies investigate the same underlying phenomenon. Small or imprecise studies can therefore appear more contradictory than their uncertainty warrants.
Do not compare only whether one study reached a conventional statistical-significance threshold and the other did not. Statistical significance is not itself an effect size, and a significant result in one study alongside a nonsignificant result in another does not automatically establish that the underlying effects differ.
Compare effect estimates, uncertainty intervals, sample sizes, design quality, and, where appropriate, formal analyses of heterogeneity or compatibility. The next replication should be designed with enough precision to help distinguish among plausible effect sizes.
Look for Implementation and Measurement Differences
A replication can follow a published methods section faithfully and still differ in implementation. Interventions may be delivered differently, laboratory equipment may behave differently, participants may interpret instructions differently, or measures may function differently across populations and contexts.
These possibilities are particularly important when tacit procedural knowledge is difficult to capture in publications. Contact with original authors, access to original materials, pilot work, manipulation checks, fidelity measures, and transparent protocols may help identify consequential differences.
But implementation differences should be investigated rather than invoked automatically. Otherwise, every failed replication becomes unfalsifiable because some unmeasured difference can always be imagined after the fact.
A Failed Replication May Reveal a Boundary Condition
Sometimes the original and replication results differ because the phenomenon genuinely depends on a condition that was not previously understood. Population characteristics, environmental conditions, institutional structures, cultural context, timing, treatment intensity, or another theoretically relevant feature may moderate the effect.
In that situation, non-replication can be scientifically productive. The National Academies distinguishes potentially helpful sources of non-replicability, which can reveal previously unrecognized variability or phenomena, from avoidable sources arising from shortcomings in research design, conduct, or communication.
The next study can therefore be more informative if it manipulates or compares the suspected condition rather than merely repeating one of the previous studies.
Sometimes the Right Next Study Is an Adjudicating Replication
When two studies disagree and several explanations remain plausible, the strongest next design may deliberately incorporate features from both.
For example, if the original and replication used different populations and different measurement instruments, a new study could cross those features where feasible. If the effect appears with one instrument but not another across comparable samples, measurement becomes a more plausible explanation. If it appears in one population under both measures but not another, population differences deserve closer attention.
This kind of design does more than count another replication. It attempts to adjudicate between explanations.
Another Near-Identical Replication Can Still Be Useful
Not every discrepancy requires a complicated factorial investigation. If both previous studies were small, another well-powered close replication may substantially improve precision. If the failed replication had an identifiable implementation problem, a corrected replication may be warranted. If the finding is highly consequential, additional independent evidence may be valuable even when the reason for disagreement remains uncertain.
The relevant standard is information gain: what will another study tell you that the two existing studies cannot?
Do Not Chase Replication Until You Get the Result You Prefer
Repeatedly conducting studies until one produces a preferred result undermines the evidential purpose of replication. Decisions about whether another study is needed, how it will be designed, and how outcomes will be interpreted should ideally be specified before the new result is known.
Preregistration or a Registered Report can be useful when a literature has become contentious because they separate evaluation of the research question and methodology from the eventual direction of the findings.
The objective is not to make the replication “succeed.” It is to reduce uncertainty.
06 · What This Means for You
Design the Next Study Around the Disagreement
If you are considering replicating a study that has already failed to replicate, begin with a side-by-side comparison of the original and replication rather than simply returning to the original paper.
Your objective is to identify what uncertainty now exists because both studies are part of the evidence base.
A simple decision framework
If both studies are reasonably rigorous but imprecise
A more precise replication may help determine which effect sizes remain compatible with the cumulative evidence.
If a specific methodological difference plausibly explains the disagreement
Design the next study to test that difference rather than merely choosing one previous procedure to copy.
If the failed replication contains a consequential implementation problem
A corrected replication may be warranted, with the problem and correction documented transparently.
If the original claim is highly consequential and the evidence remains genuinely conflicting
Additional independent evidence may have substantial value even when no single explanation for the discrepancy is yet established.
If another similar study is unlikely to distinguish among plausible explanations
Consider an adjudicating design, a targeted extension, or another research question rather than adding an uninformative repetition.
This approach also clarifies why a failed replication can still be valuable research. Its contribution may be precisely that it transforms an apparently settled result into a better scientific question about reliability, variation, or boundary conditions.
If your next study adds substantial moderators, populations, or procedures to investigate that discrepancy, however, check whether the extension still provides a meaningful test of the original finding. An elaborate explanation of disagreement is useful only if the design preserves a clear inferential connection to the claim being investigated.
07 · A Quick Checklist
Before Replicating a Study That Has Already Failed to Replicate
Before conducting another replication, check:
Read both the original study and the failed replication in full rather than relying on descriptions of the disagreement.
Compare effect estimates, uncertainty, sample sizes, and study quality rather than significance labels alone.
Map consequential differences in population, setting, materials, measures, implementation, exclusions, and analysis.
Apply comparable methodological scrutiny to the original study and the replication.
Specify plausible explanations for the disagreement before collecting the next dataset where possible.
Design the new study with enough precision to distinguish among scientifically meaningful possibilities.
Consider preregistration or a Registered Report when the literature has become contentious or interpretively flexible.
State what each plausible outcome would imply before knowing which result occurs.