01 · The Question
How Much Can You Change Before the Original Finding Is No Longer Being Tested?
A replication does not have to be a photocopy of the original study. You might want to add a moderator, recruit a new population, improve a measure, introduce another experimental condition, test a mechanism, or collect additional outcomes. These changes can make the project substantially more informative.
But there is a limit. If you change enough features at once, a result that differs from the original becomes difficult to interpret. Was the original finding unreliable? Did the effect depend on the population? Did the new measure change what was being assessed? Did the added intervention component alter the phenomenon? Or are you now investigating a different claim altogether?
The boundary between replication and extension is therefore not determined simply by counting modifications. What matters is whether the new study still contains a sufficiently clear test of the original claim.
03 · What You Need to Know
The Boundary Is Inferential, Not Numerical
Nosek and Errington propose defining replication by whether possible outcomes would provide diagnostic evidence about a claim from prior research. This shifts attention away from whether two studies look identical and toward what the new evidence can actually tell us.
That principle is particularly useful for replication-extension studies. An extension can introduce something new while preserving an interpretable test of the prior finding. Problems arise when the additions become inseparable from the test of the original claim.
Start With the Claim, Not the Original Paper as a Whole
A published study may contain several hypotheses, variables, analyses, and findings. You usually do not need to reproduce every element to test one particular claim.
Before designing an extension, state the target finding as precisely as possible. For example: “Students receiving intervention X achieved higher scores on outcome Y than students receiving the comparison condition under population and setting Z.”
Once that claim is explicit, you can ask which elements must remain sufficiently comparable for the new study to provide evidence about it.
This is also why understanding the difference between replication and extension matters before adding new components. The label should follow the inferential purpose of the study rather than being chosen merely because some procedures resemble earlier research.
Adding Something New Does Not Automatically Destroy the Replication
An extension can be layered onto a replication. Suppose the original experiment compared intervention A with control B. Your study can preserve that comparison and add condition C to investigate a new theoretical possibility.
The A-versus-B comparison may still provide a replication test, while comparisons involving C provide the extension.
Likewise, you might preserve the original primary outcome while adding secondary outcomes, or reproduce the original analysis while preregistering additional analyses. In these cases, the new material does not necessarily interfere with the original test.
Replication component
Preserves an interpretable test of the prior claim using new data.
Extension component
Asks an additional question about mechanism, generalizability, moderation, application, measurement, or another aspect beyond the original test.
The Problem Begins When the Replication Comparison Disappears
Suppose the original study compared a standard teaching method with a particular digital intervention. Your new study changes the intervention, replaces the outcome measure, recruits a substantially different population, lengthens the treatment period, and adds instructor training.
The new study may be valuable. But if it produces a different result, there are now many explanations for the discrepancy. More importantly, the design may no longer contain the comparison needed to determine whether the original effect recurs.
At that point, describing the study simply as a replication can overstate how directly it tests the original finding.
Changing the Construct Can Break the Connection Quickly
Some modifications are more consequential than they initially appear. Replacing one measure with another may seem like a technical adjustment, but if the instruments operationalize the construct differently, the substantive question may also change.
The same applies to interventions. Modifying delivery, intensity, duration, content, or participant interaction may create a meaningfully different treatment. Whether that matters depends on which features are theoretically necessary for the original claim.
Methodological improvement is therefore not automatically neutral. Even a change intended to strengthen validity can alter the target of inference.
Adding a Variable Can Change the Question Without Changing the Procedure
An extension can become conceptually different even when the original procedures remain recognizable.
Suppose the original claim is that intervention X improves learning. Your new study asks whether X improves learning only among students with high prior knowledge and whether this relationship is mediated by cognitive engagement. The original main-effect comparison may remain available, but much of the new study now concerns moderation and mechanism.
That is perfectly legitimate. The problem arises only if the researcher presents evidence about the moderator or mediator as though it were itself a replication of the original main effect.
The question of when adding a variable turns replication into extension therefore depends on the role that variable plays in the new inference.
A New Population Can Preserve or Change the Target Claim
Changing participants often moves a study toward generalization. Nature Communications has argued that replication and generalization studies can work together: close replications test robustness under similar conditions, while generalization studies investigate whether effects hold across new populations, settings, or implementations.
If the original claim was explicitly limited to one population, moving to another population tests a broader proposition. If researchers have already been treating the original result as generally applicable, however, testing a theoretically relevant new population may directly evaluate that generalization.
The issue is therefore not simply whether the sample changed. It is whether the new population provides a meaningful test of the claim.
Changing Several Things at Once Creates an Attribution Problem
Multiple simultaneous changes can make discrepant results difficult to explain. If population, measurement, setting, procedure, and analysis all differ, any one of those changes could contribute to a difference in results.
This does not automatically make the study scientifically poor. It may be an excellent test of whether a broad phenomenon appears under substantially different conditions. But it weakens your ability to attribute consistency or inconsistency to particular features of the original finding.
When identifying boundary conditions is the goal, systematic variation is generally more informative than changing several features without a design that can separate their effects.
Use the Counterfactual Interpretation Test
A practical way to evaluate your design is to imagine the results before conducting the study.
First ask: if the original pattern appears, can I reasonably say that the new evidence increases confidence in the original claim? Then ask the harder question: if the original pattern does not appear, can I reasonably say that this outcome provides evidence against, limits, or qualifies the original claim?
If the answer to both questions is unclear because too many modifications could explain either outcome, the replication component may have become too weak.
Watch Out
A design should not count as a replication only when it produces the expected result. If a positive outcome would be called confirmation but a negative outcome could always be dismissed because the extension changed too much, the study is not providing a fair diagnostic test of the original claim.
You Can Preserve Replication by Nesting the Extension Around It
One useful strategy is to preserve the original comparison as a recognizable component and build the extension around it. You might add conditions rather than replace the original conditions, add outcomes rather than discard the original outcome, or include the original population alongside a new population.
This approach can sometimes allow one study to answer two related questions: does the original finding recur, and does it extend to the new condition?
It may require a larger sample or more complex design, but the resulting evidence is often easier to interpret than a study in which every original element has been replaced simultaneously.
07 · A Quick Checklist
Before Adding More to a Replication
Before finalizing a replication-extension design, check:
State the specific original claim that the replication component is intended to test.
Identify the original comparison, condition, construct, and outcome necessary for an interpretable replication test.
Separate variables and analyses needed for replication from those introduced for the extension.
Prefer adding new conditions or outcomes alongside essential original ones when feasible rather than replacing everything at once.
Ask what a result consistent with the original finding would mean under the modified design.
Ask what a result inconsistent with the original finding would mean and whether alternative explanations created by your modifications would dominate interpretation.
Document consequential departures from the original methodology and explain why each was introduced.
Describe the project primarily as an extension when the original claim can no longer be evaluated independently of the new elements.