01 · The Question
If the literature disagrees, what would another study actually settle?
You review the literature and encounter the classic conclusion: findings are mixed.
Some studies report a positive relationship. Others report no clear association. Perhaps a few point in the opposite direction. Authors cite different populations, measures, designs, sample sizes, analytical methods, and contexts. The obvious response seems to be another study.
But imagine that you conduct essentially the same kind of study and obtain another positive result. Has the disagreement been resolved? Probably not. The literature now simply contains one additional positive result. A negative result would create the same problem in reverse.
When studies disagree, the useful next study is usually one designed to investigate why they disagree or to provide evidence capable of discriminating among the plausible explanations.
03 · What You Need to Know
How to turn conflicting findings into an informative research question
First determine whether the studies really disagree
“Study A found an effect and Study B did not” does not necessarily mean the studies provide conflicting evidence.
Researchers frequently classify studies according to whether their individual p-values cross a statistical significance threshold. This can manufacture apparent disagreement.
Imagine two studies estimating effects of similar magnitude. The larger study has a narrow confidence interval and reports p <.05. The smaller study has a wider interval and reports p >.05. The studies may be statistically compatible despite receiving different labels of “significant” and “nonsignificant.”
The first task is therefore to compare effect estimates, uncertainty intervals, study characteristics, and the substantive direction and magnitude of findings rather than simply counting statistical significance.
Different statistical significance
Studies cross a significance threshold differently, potentially because their precision differs.
Substantive inconsistency
Studies provide estimates or patterns that differ enough to affect the scientific or practical conclusion and require explanation.
Synthesize before deciding that another study is needed
When findings appear inconsistent, the first useful response may be systematic synthesis rather than immediate data collection.
Cochrane guidance treats heterogeneity as variation among study results and recommends considering whether it can be explained, including through scientifically motivated subgroup analysis or meta-regression when sufficient evidence is available. It also warns that an average effect can be misleading when effects vary importantly across studies.
Synthesis can reveal several possibilities. The estimates may actually be compatible once uncertainty is considered. There may be genuine heterogeneity. Differences may track a particular population, comparator, intervention feature, measurement approach, or design characteristic. Alternatively, the evidence may simply be too sparse to identify a convincing explanation.
Thus, before collecting another dataset, ask whether better synthesis of the existing evidence could clarify the apparent disagreement.
Identify plausible sources of disagreement
Studies can disagree for many reasons, and several can operate simultaneously.
| Possible source of disagreement |
What to examine |
What a new study might do |
| Sampling variation or imprecision |
Effect estimates and uncertainty intervals |
Provide a more precise estimate if important uncertainty remains |
| Population differences |
Characteristics plausibly related to the effect or relationship |
Test whether a specified population characteristic modifies the result |
| Measurement differences |
Definitions, instruments, thresholds, operationalizations |
Use measurement capable of testing whether operationalization explains the discrepancy
|
| Different comparators |
What each intervention, exposure, or condition was compared against |
Standardize or directly test the consequential comparison |
| Design differences |
Temporal ordering, assignment, sampling, follow-up, control of alternative explanations |
Use a design that distinguishes the competing interpretations |
| Implementation differences |
Dose, fidelity, adherence, setting, personnel, co-interventions |
Measure or manipulate the suspected implementation factor |
| Risk of bias |
Whether methodological problems differ systematically across studies |
Generate evidence less vulnerable to the consequential bias |
| Analytical choices |
Model specification, exclusions, transformations, covariate adjustment, outcome definitions |
Use prespecified analyses or compare defensible specifications where appropriate |
The table should generate hypotheses, not excuses. The goal is not to invent a post hoc explanation for every inconvenient result. Plausible explanations should have substantive or methodological support and, ideally, generate predictions that can be tested.
Population differences should become hypotheses, not labels
Suppose studies in Population A tend to report positive effects while studies in Population B do not.
It is tempting to conclude that “population explains the inconsistency.” But population labels bundle many characteristics together. The studies may also differ in measurement, intervention implementation, sampling, calendar time, or design.
A stronger next study identifies the characteristic thought to matter and tests it.
For example, if researchers hypothesize that baseline digital literacy modifies the effect of an AI learning intervention, they should measure digital literacy and design the study so that the proposed modification can be examined. Merely conducting another study in Population A or B may reproduce the pattern without explaining it.
This is the difference between simply adding another population and testing whether the population difference changes the conclusion.
Measurement differences can create apparent scientific disagreement
Two papers may use the same construct label while measuring substantially different things.
One study might define “AI use” as frequency of any generative AI interaction. Another measures use specifically for assessed coursework. A third measures perceived dependence on AI. Their findings need not agree because the variables are not equivalent.
The same problem occurs when studies use different outcome thresholds, instruments, recall periods, diagnostic criteria, coding rules, or proxies.
Before treating results as contradictory, ask whether the studies actually estimate comparable quantities.
If measurement appears to explain the disagreement, a new study can become particularly informative by measuring competing operationalizations within the same sample or by using a measurement strategy that better represents the construct. That is considerably more diagnostic than simply selecting one measure and generating another estimate.
Comparator differences can make apparently conflicting results perfectly compatible
An intervention may outperform no treatment but perform similarly to an established active intervention. Those findings are not contradictory because the comparisons answer different questions.
Cochrane emphasizes that review questions and syntheses depend on clearly defined comparisons, and heterogeneous comparisons can require separation rather than indiscriminate pooling.
If disagreement arises because studies use different comparators, the next study should focus on the comparison genuinely needed to distinguish the competing conclusions.
Risk of bias can produce differences among study results
Not every study deserves equal evidential weight simply because it exists.
Cochrane defines bias as systematic deviation from the truth in study results and notes that flaws in study design and execution can lead to overestimation or underestimation of intervention effects. Empirical meta-epidemiological evidence also indicates that some methodological characteristics are associated with systematically different effect estimates.
Suppose small studies at high risk of bias tend to report large effects while larger studies with stronger protection against bias report smaller effects. Conducting another study that shares the weaknesses of the first group is unlikely to resolve the disagreement.
A more useful study would address the methodological weakness plausibly contributing to the inconsistent evidence.
Do not confuse heterogeneity with error that must be eliminated
Sometimes studies disagree because effects genuinely vary.
Cochrane notes that heterogeneity can itself be informative when the research question concerns differential effects across populations, interventions, or circumstances.
If an intervention works differently under different conditions, forcing the literature toward one universal average may obscure the more interesting finding.
The research question can then change from “Does it work?” to “Under which conditions does it work, for whom, and to what extent?”
A new study designed around a plausible effect modifier can advance that question much more effectively than another study estimating an overall average under yet another set of conditions.
Subgroup analysis and meta-regression can suggest explanations, but caution is needed
When enough studies exist, systematic reviews may investigate whether study characteristics are associated with variation in effects. Subgroup analysis and meta-regression are common tools for this purpose.
However, Cochrane cautions that these analyses are observational comparisons across studies, may have limited power, and can generate false-positive explanations, particularly when many characteristics are explored without prior rationale.
A pattern identified in synthesis can therefore motivate a targeted new study without being treated as definitive proof of the explanation.
This is one of the strongest roles for new primary research: turn an explanatory pattern in the literature into a prospective test.
Design the new study so competing explanations make different predictions
A study becomes particularly informative when plausible explanations for the disagreement predict different outcomes.
Suppose one explanation says an intervention works only when participants receive substantial implementation support. Another says the intervention has little effect regardless of support and earlier positive findings reflect bias.
A study that measures or experimentally varies implementation support while using stronger protection against the suspected bias can provide evidence relevant to both explanations.
This is much more useful than conducting the intervention under one arbitrary level of support and then adding the result to the existing pile.
Watch Out
Do not design the new study merely to determine which “side” of the literature receives another supporting paper. Design it so that plausible explanations for the disagreement can be distinguished. Science is not improved much by turning a meta-analysis into a scoreboard.
A single new study may not resolve the disagreement completely
Some disagreements reflect several sources of heterogeneity or a literature that is simply too uncertain for one study to settle.
The appropriate standard is therefore not “Will this study end the debate?” It is whether the study will make the disagreement more interpretable.
A useful study might eliminate one explanation, support another, provide a more precise estimate under important conditions, identify a boundary condition, or reveal that apparent inconsistency was partly methodological.
That is enough to move the evidence forward without pretending that one dataset gets the final word.