Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Should You Replicate a Study That Has Already Failed to Replicate?

A failed replication does not necessarily end the scientific question. Another replication can be valuable when it can distinguish among plausible explanations for the disagreement rather than merely adding a third result to an unresolved dispute.

566
Replicating a Study After a Failed Replication Guide 566 of 603
01 · The Question

What Should You Do After an Original Finding Fails to Replicate?

The original study reports a clear finding. A later research team attempts to replicate it and obtains a substantially different result. Now the literature contains a disagreement rather than a straightforward claim.

Should you conduct another replication?

Possibly, but the rationale has changed. Before the failed replication, the question may have been whether independent evidence supports the original finding. Afterward, the more informative question is often why the studies disagree and what additional evidence could distinguish among the plausible explanations.

A third study that simply repeats one side without addressing the source of disagreement can leave the field with three results instead of two and not much more understanding.

02 · The Short Answer

Another Replication Can Be Valuable If It Helps Explain the Discrepancy

In Brief

Yes. A study that has already failed to replicate may deserve another replication when a new study can clarify whether the disagreement reflects sampling variation, methodological differences, contextual conditions, measurement, implementation, weaknesses in either study, or a genuine limit on the original claim.

Do not treat the previous failed replication as a final verdict or assume that another repetition is automatically useful. The next study should be designed around the unresolved discrepancy and should ideally distinguish among competing explanations for it.

03 · What You Need to Know

A Failed Replication Changes the Question Rather Than Ending It

The National Academies of Sciences, Engineering, and Medicine explicitly cautions that a single failed replication does not conclusively refute an original scientific claim. Non-replicability can arise for many reasons, including inherent variability, uncontrolled complexity, methodological problems, chance, or previously unrecognized phenomena.

At the same time, a failed replication should not simply be ignored. It is new evidence. The task is to determine how much that evidence should change confidence in the claim and what investigation, if any, is needed next.

“Failed to Replicate” Is Not a Single Scientific Outcome

The phrase can conceal substantial differences in what actually happened. One replication may estimate an effect near the original estimate but with wide uncertainty. Another may produce an estimate close to zero with high precision. A third may find an effect in the opposite direction. Those results do not provide identical evidence.

The National Academies notes that replicability assessment need not produce a binary pass-or-fail judgment. The degree to which results are consistent can depend on the criteria used, the uncertainty surrounding the estimates, and the scientific question being evaluated.

Before planning another study, inspect the estimates and uncertainty rather than relying only on labels such as “successful” and “failed.”

Compare the Original Study and the Failed Replication Symmetrically

Researchers sometimes assume that the original study establishes the benchmark and that the replication must explain any deviation from it. That is not necessarily justified.

Both studies are empirical observations subject to sampling variation, methodological limitations, implementation problems, and analytical choices. The original study may have overestimated an effect. The replication may have implemented an essential procedure poorly. Both could be reasonable studies conducted under conditions where the phenomenon genuinely differs.

A useful investigation therefore asks what each study contributes rather than automatically privileging whichever result appeared first.

Watch Out

Do not diagnose a failed replication solely by searching for flaws in the replication while treating the original study as error-free. Apply comparable methodological scrutiny to both.

Check Whether the Replication Tested the Same Claim Closely Enough

No two studies are literally identical. Differences in population, setting, materials, timing, measurement, implementation, and analysis are inevitable. Some differences are irrelevant to the claim; others may be consequential.

When a replication produces a discrepant result, examine whether departures from the original design could plausibly matter. This does not mean inventing a post hoc explanation for every difference. A proposed moderator or boundary condition should be scientifically defensible and, ideally, testable.

If substantial changes were introduced, you may need to revisit whether the earlier attempt was a sufficiently close replication or a more conceptual test. Understanding what kind of replication question was actually asked becomes particularly important when interpreting disagreement.

Ask Whether Sampling Variation Could Plausibly Explain the Difference

Different samples produce different estimates even when studies investigate the same underlying phenomenon. Small or imprecise studies can therefore appear more contradictory than their uncertainty warrants.

Do not compare only whether one study reached a conventional statistical-significance threshold and the other did not. Statistical significance is not itself an effect size, and a significant result in one study alongside a nonsignificant result in another does not automatically establish that the underlying effects differ.

Compare effect estimates, uncertainty intervals, sample sizes, design quality, and, where appropriate, formal analyses of heterogeneity or compatibility. The next replication should be designed with enough precision to help distinguish among plausible effect sizes.

Look for Implementation and Measurement Differences

A replication can follow a published methods section faithfully and still differ in implementation. Interventions may be delivered differently, laboratory equipment may behave differently, participants may interpret instructions differently, or measures may function differently across populations and contexts.

These possibilities are particularly important when tacit procedural knowledge is difficult to capture in publications. Contact with original authors, access to original materials, pilot work, manipulation checks, fidelity measures, and transparent protocols may help identify consequential differences.

But implementation differences should be investigated rather than invoked automatically. Otherwise, every failed replication becomes unfalsifiable because some unmeasured difference can always be imagined after the fact.

A Failed Replication May Reveal a Boundary Condition

Sometimes the original and replication results differ because the phenomenon genuinely depends on a condition that was not previously understood. Population characteristics, environmental conditions, institutional structures, cultural context, timing, treatment intensity, or another theoretically relevant feature may moderate the effect.

In that situation, non-replication can be scientifically productive. The National Academies distinguishes potentially helpful sources of non-replicability, which can reveal previously unrecognized variability or phenomena, from avoidable sources arising from shortcomings in research design, conduct, or communication.

The next study can therefore be more informative if it manipulates or compares the suspected condition rather than merely repeating one of the previous studies.

Sometimes the Right Next Study Is an Adjudicating Replication

When two studies disagree and several explanations remain plausible, the strongest next design may deliberately incorporate features from both.

For example, if the original and replication used different populations and different measurement instruments, a new study could cross those features where feasible. If the effect appears with one instrument but not another across comparable samples, measurement becomes a more plausible explanation. If it appears in one population under both measures but not another, population differences deserve closer attention.

This kind of design does more than count another replication. It attempts to adjudicate between explanations.

Another Near-Identical Replication Can Still Be Useful

Not every discrepancy requires a complicated factorial investigation. If both previous studies were small, another well-powered close replication may substantially improve precision. If the failed replication had an identifiable implementation problem, a corrected replication may be warranted. If the finding is highly consequential, additional independent evidence may be valuable even when the reason for disagreement remains uncertain.

The relevant standard is information gain: what will another study tell you that the two existing studies cannot?

Do Not Chase Replication Until You Get the Result You Prefer

Repeatedly conducting studies until one produces a preferred result undermines the evidential purpose of replication. Decisions about whether another study is needed, how it will be designed, and how outcomes will be interpreted should ideally be specified before the new result is known.

Preregistration or a Registered Report can be useful when a literature has become contentious because they separate evaluation of the research question and methodology from the eventual direction of the findings.

The objective is not to make the replication “succeed.” It is to reduce uncertainty.

04 · A Practical Example

Two Studies Disagree About an Educational Intervention

Hypothetical Example

A Strong Original Effect Is Followed by a Near-Zero Replication

Suppose an original experiment reports that a digital feedback intervention substantially improves student performance. An independent replication later reports an effect close to zero. Both studies appear reasonably conducted, but they differ in several potentially relevant ways.

Compare the evidence The researchers examine effect estimates and uncertainty rather than merely noting that the first study was statistically significant and the second was not.
Map consequential differences The original study used first-year university students and weekly feedback over an entire semester. The replication used senior students and a shorter intervention period.
Develop competing explanations The original estimate may have been unusually large, the intervention duration may matter, student level may moderate the effect, or some combination of these possibilities may explain the discrepancy.
Design the next replication Rather than simply copying one study, the new investigation uses adequate precision and deliberately tests whether intervention duration or student level changes the effect while preserving the intervention's essential features.
Interpret the contribution The study can potentially explain why previous results diverged and refine the conditions under which the original claim should be expected to hold.

At this stage, the scientific contribution is no longer merely “we replicated the original study again.” The stronger contribution is resolving an evidential conflict.

05 · What Researchers Often Get Wrong

Common Mistakes After a Failed Replication

Misconception

The Original Finding Has Been Disproved

A failed replication is evidence against straightforward confidence in the original result, but one study does not usually provide a conclusive verdict. Interpretation should consider the quality, precision, comparability, and cumulative evidence from both studies.

Misconception

The Replication Must Have Been Done Incorrectly

That is one possibility, not a default explanation. The original study can also contain limitations, and genuine variation between contexts may exist. Scrutinize both studies using comparable standards.

Misconception

A Third Study Will Automatically Settle the Disagreement

Another study helps only to the extent that its design and precision add diagnostic evidence. Repeating an ambiguous design can simply produce another ambiguous result.

Misconception

Significant in the Original and Nonsignificant in the Replication Means the Results Conflict

Different significance classifications do not by themselves demonstrate that the effect estimates differ. Compare estimates and their uncertainty and use an appropriate inferential framework for the substantive question.

Misconception

Any Difference Between the Studies Explains the Failed Replication

After a discrepancy appears, it is easy to generate explanations involving population, culture, timing, materials, or implementation. A plausible explanation becomes scientifically useful when it is specified clearly enough to be tested rather than merely invoked after the fact.

Misconception

A Failed Replication Has No Value Unless Someone Resolves It

The failed replication already contributes evidence by revealing that the original result did not recur under the replication conditions. Further work may clarify why, but the discrepant result itself can still be informative.

06 · What This Means for You

Design the Next Study Around the Disagreement

If you are considering replicating a study that has already failed to replicate, begin with a side-by-side comparison of the original and replication rather than simply returning to the original paper.

Your objective is to identify what uncertainty now exists because both studies are part of the evidence base.

A simple decision framework

If both studies are reasonably rigorous but imprecise
A more precise replication may help determine which effect sizes remain compatible with the cumulative evidence.
If a specific methodological difference plausibly explains the disagreement
Design the next study to test that difference rather than merely choosing one previous procedure to copy.
If the failed replication contains a consequential implementation problem
A corrected replication may be warranted, with the problem and correction documented transparently.
If the original claim is highly consequential and the evidence remains genuinely conflicting
Additional independent evidence may have substantial value even when no single explanation for the discrepancy is yet established.
If another similar study is unlikely to distinguish among plausible explanations
Consider an adjudicating design, a targeted extension, or another research question rather than adding an uninformative repetition.

This approach also clarifies why a failed replication can still be valuable research. Its contribution may be precisely that it transforms an apparently settled result into a better scientific question about reliability, variation, or boundary conditions.

If your next study adds substantial moderators, populations, or procedures to investigate that discrepancy, however, check whether the extension still provides a meaningful test of the original finding. An elaborate explanation of disagreement is useful only if the design preserves a clear inferential connection to the claim being investigated.

07 · A Quick Checklist

Before Replicating a Study That Has Already Failed to Replicate

Before conducting another replication, check:
Read both the original study and the failed replication in full rather than relying on descriptions of the disagreement.
Compare effect estimates, uncertainty, sample sizes, and study quality rather than significance labels alone.
Map consequential differences in population, setting, materials, measures, implementation, exclusions, and analysis.
Apply comparable methodological scrutiny to the original study and the replication.
Specify plausible explanations for the disagreement before collecting the next dataset where possible.
Design the new study with enough precision to distinguish among scientifically meaningful possibilities.
Consider preregistration or a Registered Report when the literature has become contentious or interpretively flexible.
State what each plausible outcome would imply before knowing which result occurs.
08 · Frequently Asked Questions

Questions About Replicating After a Failed Replication

How many failed replications does it take to reject a finding?

There is no universal number. Scientific conclusions should reflect the cumulative quality, precision, independence, comparability, and consistency of the evidence rather than a simple count of successful and unsuccessful replications.

Should the next study copy the original or the failed replication?

Neither automatically. First identify what the disagreement leaves unresolved. The most informative design may resemble one study closely, correct an identifiable problem, or deliberately combine features of both to test competing explanations.

Does a nonsignificant replication count as a failed replication?

Not necessarily. A nonsignificant result can be highly imprecise and compatible with meaningful effects. Replication assessment should consider effect estimates, uncertainty, design, and the criterion used to evaluate consistency rather than relying only on a significance threshold.

What if the replication found a much smaller effect rather than no effect?

That can be scientifically important. The evidence may suggest that the original effect was overestimated while still supporting a smaller phenomenon. Compare estimates and uncertainty rather than forcing the studies into replicated versus not replicated categories.

Can different populations explain a failed replication?

They can if a relevant population characteristic plausibly moderates the phenomenon. The explanation becomes stronger when specified in advance and tested directly rather than inferred solely because two studies happened to use different samples.

Is another replication worthwhile if the original study had serious methodological weaknesses?

Potentially. If the claim remains important, another study can be useful, but simply repeating a design that cannot distinguish the target explanation from alternatives may not resolve the disagreement. A stronger test may be preferable.

Can a second failed replication still be publishable?

Scientific value should not depend on obtaining a particular direction of result. A rigorous study can contribute useful evidence by improving precision, testing an explanation for previous disagreement, or clarifying the conditions under which a claim does or does not hold.

09 · The Bottom Line

After a Failed Replication, Replicate the Uncertainty Rather Than the Argument

The Bottom Line

A study that has already failed to replicate can deserve another replication when the next study is capable of clarifying why the evidence conflicts or of producing substantially more informative evidence about the original claim.

Treat neither the original study nor the failed replication as the automatic winner. Compare both critically, identify what remains unresolved, and design the next investigation so that it reduces uncertainty rather than merely adding another result to the dispute.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes