03 · What You Need to Know
How to determine whether independent replication would add useful evidence
Replication is not simply doing the same study again
A common definition of replication is straightforward: repeat the original procedures and see whether the result occurs again. That captures an important practical form of replication, but it can obscure the scientific purpose.
Nosek and Errington propose a more inferential definition: a replication is a study for which the possible outcomes provide diagnostic evidence about a claim from prior research. Under this view, evidence consistent with the prior claim should increase confidence in it, while inconsistent evidence should decrease confidence or force the claim's boundaries to be reconsidered.
This matters because no replication reproduces every condition of the original study exactly. Participants differ. Time passes. Researchers differ. Equipment, settings, implementation details, and countless incidental circumstances may change.
The scientific task is to decide which conditions must remain sufficiently similar for the new evidence to test the same claim and which conditions should be irrelevant if the claim is correct.
Replication and reproducibility are not the same problem
The terminology varies somewhat among fields, but an influential distinction from the U.S. National Academies defines replicability as obtaining consistent results across studies addressing the same scientific question using new data. Reproducibility, in that framework, concerns obtaining consistent computational results using the same data, computational procedures, code, and conditions of analysis.
Reproducibility
Can the reported computational result be obtained again from the original data using the specified analytical procedures?
Replicability
Does new data collected or obtained in another study provide results consistent with the scientific claim under conditions where consistency is expected?
Both matter, but they diagnose different problems. Re-running code may reveal a computational error. It does not establish that the phenomenon will appear in another sample. Conversely, a successful replication with new data does not demonstrate that the original analysis itself was computationally reproducible.
Independence matters because it changes what is being tested
Replication becomes especially informative when the new evidence is not merely another product of the same dependencies that generated the original result.
Independence can involve new participants or observational units, a new dataset, different investigators, independent implementation, or other separation from the original research process. The exact form depends on the field.
Why does this matter? Suppose the original result depended unknowingly on a laboratory-specific procedure, an idiosyncratic coding decision, a local recruitment process, or an undocumented analytical convention. A study conducted by the same team using nearly identical infrastructure may reproduce that dependency along with the intended phenomenon.
An independent implementation can test whether the claim survives beyond those conditions.
Independence should not be romanticized, however. A completely independent team can misunderstand the original method. Independence increases the value of the test only when the replication is competent and sufficiently aligned with the claim being evaluated.
Replication is particularly valuable when confidence rests heavily on one influential result
The marginal value of replication depends partly on the existing evidence.
If a consequential claim rests largely on one study, independent evidence may substantially change confidence in it. If numerous rigorous and genuinely independent studies already test the same claim under relevant conditions, another nearly identical replication may contribute less.
This is why replication should be evaluated against the evidence base rather than the fame of one paper.
Before collecting new data, determine whether independent evidence already exists and whether better synthesis could establish how much replication has already occurred.
A useful replication makes both broad outcomes informative
Imagine proposing a replication and deciding in advance how you would interpret its possible results.
If a result consistent with the original finding would be presented as confirmation, what would happen if the result were inconsistent?
If every inconsistent result could immediately be dismissed because the population, year, researcher, setting, or minor procedural detail differed, then the study may not provide a strong test of the original claim. It may instead be testing generalizability to a new condition.
Nosek and Errington emphasize this symmetry: a replication is most clearly diagnostic when both consistency and inconsistency would update confidence in the prior claim.
Watch Out
Do not define the study as a replication only after seeing whether the result agrees with the original. Decide beforehand which differences from the original study are compatible with testing the same claim and how both consistent and inconsistent outcomes would be interpreted.
Close procedural similarity can be valuable when theory is underspecified
Researchers sometimes distinguish “direct” or “close” replication from “conceptual” replication. The terminology is debated, but the practical distinction captures a real design choice.
When researchers do not yet understand which methodological conditions are essential to producing the phenomenon, staying close to the original procedures reduces the number of differences that could explain a discrepant result. Nosek and Errington note that procedural similarity can therefore be especially useful when theoretical and methodological understanding is immature.
Suppose an original experiment uses a particular manipulation, outcome measure, timing protocol, and participant population. Changing all four simultaneously may produce an interesting test, but if the result differs, identifying why becomes difficult.
A close replication deliberately limits those interpretive degrees of freedom.
Changing conditions can test whether the claim has broader scope
Replication does not require freezing every feature of the original study.
Every new study inevitably changes some conditions, and those differences can become scientifically useful. If a claim is expected to hold despite a change in researcher, setting, equipment, population, or implementation, observing it again provides evidence that the finding is not restricted to the exact historical circumstances of the original study.
But this requires theoretical clarity. If the new population is deliberately chosen because researchers expect the effect might differ, the study may be better characterized as testing a boundary condition or generalizability rather than simply replication.
That distinction becomes particularly important when the proposed contribution is a population that could meaningfully change what can be concluded.
A replication should be capable of detecting evidence that matters
Replication is not informative merely because it exists.
A study with very low precision may produce a result compatible with both the original claim and important alternatives. Likewise, poor measurement, weak implementation, severe attrition, or inappropriate analysis can make a discrepant result difficult to interpret.
The replication should therefore be designed so that its expected precision and methodological quality are adequate for the claim being tested.
This does not mean that a replication must reproduce the original effect estimate exactly. Sampling variation alone means estimates will differ. The relevant interpretation should consider effect estimates and their uncertainty, not merely whether one study reports p <.05 and another does not.
A successful replication does not prove that the original claim is universally true
Replication contributes evidence cumulatively.
A consistent result increases confidence under the conditions represented by the studies. It does not demonstrate that the claim holds for every population, setting, implementation, measurement strategy, or historical period.
Likewise, an inconsistent result does not automatically prove that the original study was false. The difference could reflect sampling variability, methodological problems, previously unidentified boundary conditions, or other consequential differences between studies.
This is why replication results should be interpreted as evidence that updates a body of knowledge rather than as binary verdicts on individual papers.
Replication becomes stronger when design decisions are transparent
Replication is particularly vulnerable to hindsight because researchers already know the original result. Analytical and methodological flexibility can therefore affect how convincing the new evidence appears.
Prospective specification of hypotheses, outcomes, exclusions, analyses, and criteria for interpreting the replication can reduce ambiguity about which decisions were made before versus after results were known.
Open materials, code, and sufficiently detailed methods can also help other researchers understand differences between the original and replication studies.
The goal is not bureaucratic perfection. It is to make the evidential confrontation between the prior claim and the new data as interpretable as possible.
Replication can be valuable even when the result does not replicate
A discrepant replication is not necessarily a failed study.
If the study was capable of providing diagnostic evidence, inconsistency can reveal that confidence in the original claim should decrease, that the effect is less stable than previously thought, or that an unrecognized condition influences when the phenomenon appears.
Nosek and Errington argue that repeated tests of replicability can help clarify the conditions under which evidence is expected to recur, thereby refining theoretical claims.
The value of replication therefore does not depend on obtaining the “right” answer. It depends on whether either answer teaches us something about the claim.