01 · The Question
Which Findings Keep Appearing When Researchers Test Them Again?
A striking result from one study can attract attention, but scientific confidence usually depends on what happens when the underlying question is investigated again. Does the finding recur with new data? Does it survive investigation by independent researchers? Does it remain visible in different populations, settings, measurements, or methodological implementations?
Answering those questions is more complicated than counting papers that report similar conclusions. Multiple publications may use the same dataset. Several studies may come from the same research group. Researchers may use “replication” differently across disciplines. And results need not be numerically identical to be scientifically consistent.
The task is therefore to map which findings recur across genuinely informative studies, how independent those studies are, what changed between them, and how much consistency the evidence actually shows.
03 · What You Need to Know
How to Determine Which Findings Have Really Been Replicated
Define Replication Before You Count It
Terminology varies across scientific disciplines, so state the definition you are using. The National Academies distinguishes computational reproducibility from replicability. Under its terminology, reproducibility means obtaining consistent computational results using the same input data, code, methods, and analytical conditions, while replicability means obtaining consistent results across studies addressing the same scientific question with new data.
Reproducibility
Under the National Academies definition, the original data and computational procedures are used to determine whether the reported computational results can be obtained consistently.
Replicability
New data are used in studies addressing the same scientific question to determine whether results are sufficiently consistent.
Other disciplines and organizations sometimes reverse these terms or use them differently. The National Academies explicitly notes this terminological inconsistency.
For literature mapping, define your terminology once and apply it consistently rather than assuming every paper uses “replication” in the same sense.
Identify the Finding, Not Merely the Paper
A paper can contain many findings. Before asking whether something replicated, specify the claim being evaluated.
For example:
- Variable A is positively associated with variable B;
- intervention X improves outcome Y relative to comparison Z;
- effect X is larger under condition A than condition B;
- a proposed mechanism mediates a relationship; or
- a qualitative pattern recurs across comparable contexts.
The claim should be specific enough that you can determine whether subsequent evidence addresses substantially the same scientific question.
Do Not Count Multiple Papers From the Same Data as Independent Replications
Several articles can arise from one dataset, cohort, trial, survey, or experiment. They may analyze different outcomes or models, but they do not constitute independent new-data replications of the same finding merely because they appear as separate publications.
Track datasets and samples where possible. Look for overlapping recruitment dates, sample sizes, project names, trial registrations, cohort names, institutions, author descriptions, or explicit references to earlier datasets.
This becomes particularly important when assessing findings that may depend heavily on one study or research group.
Researcher Independence Adds Information, but Is Not the Definition of Replication
A replication can be conducted by the original researchers or by another team. The National Academies' definition focuses on new data and the same scientific question rather than requiring different investigators.
Nevertheless, independent research groups can provide additional evidential value because they are less likely to reproduce every laboratory-specific procedure, analytic habit, recruitment channel, or unrecognized assumption of the original team.
When mapping replication, therefore, record both whether new data were collected and whether investigators, laboratories, datasets, settings, or other relevant features are independent.
Distinguish Close Replications From Tests Under Changed Conditions
Some replication studies attempt to reproduce the original design and conditions as closely as practical. Others test the same or related claim using different operationalizations, contexts, populations, or methods.
Terminology for these approaches varies, but the distinction is useful. A close replication asks whether the result recurs under conditions resembling the original study. A broader or conceptual replication can ask whether the underlying claim survives meaningful changes in how it is tested. Nature Human Behaviour has discussed direct and conceptual replication as complementary approaches to examining robustness.
| Evidence pattern |
What it can tell you |
Important question |
| New data, closely similar design |
Whether the original result recurs under similar conditions |
Were the procedures sufficiently comparable? |
| New data, different population or setting |
Whether the finding extends beyond the original context |
Are differences theoretically relevant? |
| New data, different operationalization or method |
Whether the underlying claim survives a different test |
Are both studies genuinely testing the same construct or proposition? |
| Same data, repeated analysis |
Whether computational or analytical results can be reproduced or are robust to alternative analysis |
Is this being confused with new-data replication? |
| Multiple papers using overlapping data |
Additional analyses of substantially the same evidence |
How independent are the apparent confirmations? |
Do Not Define Successful Replication as “Both p <.05”
Two studies can estimate similar effects while one crosses an arbitrary significance threshold and the other does not. Conversely, two large studies can both report statistically significant results while estimating meaningfully different effect sizes.
The National Academies describes replicability in terms of consistency given the uncertainty inherent in the system being studied and notes that there is no single universal criterion for determining replication across all disciplines.
Depending on the research area, relevant considerations may include:
- direction of effect;
- effect magnitude;
- confidence or uncertainty intervals;
- compatibility of estimates;
- substantive rather than merely statistical significance;
- prespecified replication criteria;
- measurement precision; and
- known contextual differences between studies.
The appropriate criterion should reflect the scientific claim, not merely whether two p-values fall on the same side of a threshold.
Consistency Does Not Mean Identical Results
New samples naturally produce different estimates. Populations, implementation, measurement, and random variation also contribute differences.
If the original study estimates an effect of 0.30 and a later study estimates 0.24, that is not automatically a replication failure. Nor does an estimate of 0.31 automatically establish replication if the studies differ fundamentally in what they measure.
Ask whether the results are sufficiently compatible to support the same substantive claim given their uncertainty and design.
Repeated Findings Across Changed Conditions Can Reveal Boundaries
Replication is informative even when findings do not recur uniformly.
Suppose an effect appears repeatedly among adults but not adolescents, or under controlled conditions but not routine implementation. That pattern may reveal a boundary condition rather than simply a binary success or failure.
The National Academies distinguishes replicability from generalizability, with generalizability concerning whether results apply in different contexts or populations.
This is why replication mapping should preserve differences among populations, contexts, and methods rather than simply tallying “successful” and “failed” studies.
Replication and Evidence Strength Are Related but Not Identical
Repeated independent findings can increase confidence in a scientific claim, and replication is one important way scientific knowledge accumulates. The National Academies describes replication as a key mechanism for building confidence in scientific results.
But replicated findings can still have limitations. Studies might repeatedly use biased measurements, narrow populations, or designs incapable of supporting the causal interpretation being attached to the finding.
Replication should therefore contribute to, rather than replace, your assessment of where the evidence is strong.
One Non-Replication Does Not Automatically Erase the Original Finding
The reverse is also important. A single replication attempt that produces a different result does not necessarily establish that the original finding was false. Differences in precision, implementation, population, measurement, contextual conditions, or chance may matter.
The National Academies cautions that one successful replication does not prove an original result correct and one unsuccessful replication does not conclusively refute it.
The appropriate response is to examine the total pattern of evidence and determine why studies may differ.
Map Replication as a Network of Evidence
Instead of classifying each paper in isolation, consider building a finding-level map.
For each major finding, record the originating study, subsequent studies testing the same claim, whether they use new data, researcher or dataset independence, methodological similarity, population and setting, effect estimates or substantive conclusions, and whether the result is consistent with the original claim.
This reveals whether a finding rests on one influential paper, several closely related studies, or a genuinely distributed body of independent evidence.