03 · What You Need to Know
Replication Can Test More Than Whether a Result Appears Twice
Replication and generalizability ask related but different questions
Replication concerns whether new evidence provides an informative test of an existing claim. Generalizability concerns how far the resulting inference can extend beyond the populations, settings, procedures, or circumstances already studied.
Nosek and Errington propose that replication is best understood through the outcomes of a study: a study is a replication when its possible results would provide diagnostic evidence about a claim from previous research. This framing is useful for local replication because it shifts attention away from merely copying procedures.
The question becomes: what would repeating this study here tell us about the original claim?
If the new context introduces a relevant variation, the replication can help establish whether the claim survives that variation. If the context is nearly interchangeable with settings already represented, another replication may strengthen confidence in the finding but contribute relatively little new evidence about its generalizability.
Replication evidence
New evidence tests whether a previous claim is supported again.
Generalizability evidence
New evidence helps determine whether and under what conditions that claim extends to other populations, settings, or circumstances.
Generalizability is about the scope of an inference
A study does not simply possess or lack external validity in the abstract. Generalizability is always relative to a target population, setting, intervention, outcome, or set of conditions.
A finding established among university students may generalize to students at other universities while remaining uncertain for working adults outside higher education. An educational intervention may generalize across schools with similar resources while remaining uncertain in schools where the delivery mechanism cannot be implemented as intended.
Accordingly, a useful local replication begins by defining the target of generalization. To whom, where, and under what circumstances is the original claim supposed to apply?
Choose the new context because of what it tests
The strongest local replication is not necessarily conducted in the most geographically distant location. It is conducted in a context that provides informative variation.
Suppose previous studies of an intervention were conducted in affluent urban schools. A replication in another affluent urban school system on another continent certainly adds geographical diversity. A replication in resource-constrained rural schools might provide a more demanding test if resources, staffing, class size, or infrastructure are plausibly connected to how the intervention works.
This is why the scientific importance of contextual differences matters more than distance on a map.
Contextual variation should be connected to a mechanism
Replication becomes more informative when researchers can explain why the new condition might matter.
If an intervention works partly because teachers receive extensive training, then replicating it where teachers receive much less training tests something relevant to the mechanism. If a behavioral relationship is expected to depend on social norms, a population with meaningfully different norms may provide an informative test. If an instrument assumes particular linguistic distinctions, using it in another language raises questions about measurement equivalence.
Without such reasoning, an inconsistent result can become difficult to interpret. Was the original claim limited? Was the intervention implemented differently? Did the measurement change? Was the study underpowered? Or did ordinary sampling variation produce the discrepancy?
A contextual hypothesis does not eliminate these possibilities, but it makes the replication more diagnostic.
Replication requires enough comparability to interpret similarity and difference
There is a tension in local replication. If everything is kept identical, the study may tell us little about generalizability across changing conditions. If too many things change simultaneously, a different result becomes difficult to attribute to any particular factor.
The design therefore needs a defensible balance.
Core constructs, outcomes, and analytical logic should usually remain sufficiently comparable to the original evidence. At the same time, adaptations may be necessary to preserve the substantive meaning of the study in a different context. Literal procedural sameness is not always equivalent to conceptual sameness.
For example, translating a questionnaire word for word may preserve surface form while changing meaning. Conversely, carefully adapting wording while validating the construct may create a more meaningful comparison.
A successful local replication can expand the supported scope of a finding
Suppose an effect has been demonstrated repeatedly under one narrow set of conditions. A well-designed replication finds a comparable effect in a substantially different population where there was a credible reason to question whether the mechanism would operate similarly.
That result does more than produce another positive finding. It provides evidence that the claim is robust to the particular contextual difference examined.
Notice the qualification: the replication supports generalization across the tested variation. It does not demonstrate universality. Evidence from two countries does not establish that a finding applies to every country, just as evidence from two universities does not establish that it applies to all universities.
An unsuccessful replication can reveal a boundary condition
A different result can also be scientifically valuable. If an effect appears reliably under one set of conditions but weakens or disappears under another, the discrepancy may identify a boundary condition: a circumstance under which the original claim changes or ceases to hold.
However, one failed replication does not automatically prove a contextual boundary. Differences can result from sampling error, measurement, implementation, analytical decisions, statistical power, or biases in either the original or replication study.
The useful question is not “Which study is correct?” but “What pattern of evidence best explains why the results differ?”
Watch Out
Do not interpret a significant result in one country and a non-significant result in another as evidence that the countries differ. A difference between statistical significance levels is not itself evidence of a statistically meaningful difference between effects. Directly estimate and compare the relevant effects, with uncertainty, using an appropriate design or synthesis.
Effect sizes and uncertainty matter more than matching p-values
A common replication mistake is to classify studies simply as “successful” when both produce statistically significant results and “failed” when the replication does not.
That approach discards important information. A replication may estimate an effect similar to the original but with wider uncertainty because its sample is smaller. Conversely, both studies may produce statistically significant results while estimating meaningfully different effects.
Compare effect estimates, confidence intervals or other appropriate uncertainty measures, and the substantive magnitude of differences. Where several replications exist, evidence synthesis may provide a more informative picture of heterogeneity across settings than pairwise declarations of success or failure.
Local replication becomes especially useful when previous evidence is contextually narrow
If an existing claim rests predominantly on one country, demographic group, institutional system, language, or resource environment, evidence from a meaningfully different setting can substantially broaden the empirical base.
The contribution may be smaller when previous research already spans the relevant variation. If an intervention has been evaluated successfully across many countries, resource environments, populations, and institutional arrangements, another highly similar local study may have limited incremental value.
Before proposing replication, examine the distribution of existing evidence, not merely whether your particular locality appears in the literature.
Local replication can be designed prospectively as a test of heterogeneity
A particularly strong design specifies before data collection why a contextual factor might alter the effect. Researchers can measure that factor explicitly, preregister relevant hypotheses or analyses where appropriate, and ensure sufficient precision to examine the anticipated difference.
Multisite research can be even more informative because variation across contexts is built into the design rather than inferred after comparing isolated studies. When feasible, coordinated protocols across sites can help separate contextual heterogeneity from procedural differences.
Local replication does not have to mean one local team working in isolation.
The goal is cumulative evidence, not a contest between countries
Replication is sometimes narrated as though the new study must either vindicate or overturn the original. That framing is unnecessarily adversarial and scientifically limiting.
Individual studies are estimates produced under particular conditions. A local replication contributes another piece of evidence. Similarity can increase confidence in robustness across the tested settings. Difference can motivate examination of moderators, mechanisms, measurement, implementation, or bias.
Either way, the objective is a more accurate account of where the phenomenon holds and why.
Local relevance and broader contribution can reinforce each other
A replication may be motivated partly by a local decision. A healthcare system may want to know whether an intervention supported internationally works under its own staffing model. A school system may need evidence about an instructional program under local class sizes and technology constraints.
When designed carefully, such research can serve the local decision while also contributing to broader understanding of generalizability. This is one route through which a local study can contribute beyond its immediate setting.
The key is to make the contextual features explicit enough that researchers elsewhere can understand what the local result actually tests.