01 · The Question
When Studies Differ, Is Another Study Really the Answer?
You review the existing research and find substantial variation. Effects are large in some studies and small in others. Perhaps the direction of the association changes across settings, populations, or methods. The literature clearly does not tell one tidy story.
That heterogeneity can look like an obvious research gap: results vary, therefore another study is needed. Yet another study may simply add one more estimate to the variation you already have.
The more useful question is whether new primary data can help explain, distinguish, or reduce an important source of uncertainty that the existing evidence cannot resolve.
03 · What You Need to Know
Not All Heterogeneity Creates the Same Research Need
Heterogeneity broadly refers to variability among studies. Cochrane distinguishes several forms. Differences in participants, interventions, and outcomes can be described as clinical diversity, while differences in study design, measurement tools, and risk of bias constitute methodological diversity. Statistical heterogeneity refers to variation in observed effects beyond what might reasonably be attributed to sampling variation alone.
Although this terminology developed largely within systematic reviews of interventions, the underlying distinction is useful much more broadly. Studies can disagree because the phenomenon genuinely varies, because researchers studied different versions of it, because their methods differ, or simply because estimates contain sampling uncertainty.
Meaningful heterogeneity
Variation that may reflect genuine differences in effects or relationships across populations, settings, interventions, exposures, or other theoretically relevant conditions.
Methodological heterogeneity
Variation associated with differences in study design, measurement, analysis, implementation, or risk of bias rather than necessarily with the underlying phenomenon itself.
Those possibilities lead to different next studies. If you have not established which kind of variation matters, simply adding another primary study is premature.
First Ask What Might Be Producing the Variation
Suppose studies of the same educational intervention report effects ranging from negligible to substantial. One interpretation is that the intervention genuinely works better in some circumstances. Another is that the studies implement different versions of the intervention. Perhaps outcomes are measured differently, follow-up periods vary, or stronger effects tend to occur in studies with greater risk of bias.
Those explanations are scientifically different even though they all appear initially as “heterogeneous findings.”
A systematic review can help map those differences. Where enough suitable studies exist, subgroup analyses or meta-regression may sometimes investigate whether study characteristics are associated with effect variation. Such analyses require caution, particularly when few studies are available or when explanations are developed after inspecting the results.
Another Primary Study Is Valuable When Existing Studies Cannot Test the Suspected Explanation
Suppose synthesis suggests that effects may differ according to intervention intensity, but the existing studies use inconsistent definitions of intensity and provide insufficient information for a credible comparison. A new primary study could deliberately manipulate or compare intensity under controlled conditions.
Now the new study has a specific job. It is not merely adding another effect estimate. It is testing an explanation for why previous effect estimates differ.
The same logic applies when variation appears related to age, setting, implementation fidelity, exposure level, baseline characteristics, measurement approach, or another plausible effect modifier. If existing evidence cannot distinguish among those possibilities, purpose-built primary research may be informative.
Heterogeneity Can Reveal a Missing Population or Context
Sometimes existing variation suggests that findings are context dependent, while an important context remains poorly represented.
Imagine that studies from highly resourced universities tend to report stronger effects from a digital learning intervention than studies from institutions with limited technological infrastructure. If very few studies represent the latter environments, another generic university study would contribute little. A well-designed study in the underrepresented context could be much more valuable.
The justification comes from what the study adds to the evidence structure, not merely from conducting the research somewhere new.
Do Not Treat I2 as a Traffic Light for New Research
In meta-analysis, I2 is commonly used to describe the proportion of observed variation associated with heterogeneity rather than sampling error. Cochrane provides rough interpretive ranges but explicitly cautions that the importance of an I2 value depends on factors such as the magnitude and direction of effects and the strength of evidence for heterogeneity. Uncertainty can be substantial when few studies are available.
An I2 of 70%, for example, does not translate into “70% heterogeneity” in a simple substantive sense, nor does it automatically mean that another primary study is required.
The number is a diagnostic aid. The research decision still depends on what is varying, why that variation matters, and whether new data can clarify it.
A Random-Effects Model Does Not Make Heterogeneity Go Away
A random-effects meta-analysis assumes that studies estimate different but related underlying effects and summarizes their distribution under specified assumptions. It can therefore incorporate between-study variation into the statistical model.
But Cochrane explicitly cautions that using a random-effects model is not a substitute for investigating heterogeneity. A pooled average may be particularly incomplete when effects vary considerably across contexts. Even a precise estimate of the average effect does not tell you that the effect will be similar everywhere.
This is why deciding whether meta-analysis can answer the question better than another primary study requires attention to the nature of the variation rather than merely whether a pooled estimate can technically be calculated.
Sometimes the Correct Response Is Better Synthesis, Not More Data
If many relevant studies exist but their differences have never been examined systematically, the immediate problem may still be synthesis. Collecting new data before understanding the structure of existing variation risks adding another study whose meaning becomes clear only after someone eventually conducts the review.
This is especially important when deciding whether to collect new data when existing studies have never been synthesized properly. A synthesis can reveal whether heterogeneity follows recognizable patterns and which uncertainties genuinely remain.
Sometimes the Studies Are Simply Too Different
Not every collection of studies represents one underlying research question. Studies may use such different constructs, interventions, outcomes, populations, or designs that describing them as heterogeneous understates the problem.
If studies are not meaningfully comparable, calculating an overall effect may obscure rather than clarify the evidence. Cochrane notes that a systematic review need not contain a meta-analysis and that statistical pooling may be misleading when results vary considerably, particularly when effects differ in direction.
Determining when a lack of comparable studies makes synthesis impossible or severely limited should therefore precede any conclusion that statistical heterogeneity itself demands new data.
Design the New Study Around the Source of Uncertainty
Once heterogeneous evidence genuinely supports further primary research, the design should follow from the suspected source of variation.
If population differences matter, sample the relevant populations. If implementation varies, standardize or deliberately compare implementation. If measurement differences obscure interpretation, use stronger or harmonized measures. If study quality appears associated with results, design a study that addresses the recurring methodological weakness.
“Previous studies produced heterogeneous findings” is therefore the beginning of a justification, not the completed justification.
Watch Out
Do not design another ordinary study and then cite heterogeneity as its rationale. Specify what feature of the heterogeneity remains unexplained, why resolving it matters, and how your design can provide information that the existing studies cannot.
07 · A Quick Checklist
Before Using Heterogeneity to Justify Another Study
Before collecting new data, check:
Identify whether the observed variation is clinical or substantive, methodological, statistical, or some combination of these.
Examine the magnitude and direction of study findings rather than interpreting a heterogeneity statistic in isolation.
Determine whether existing studies can already investigate plausible explanations for the variation.
Treat post hoc subgroup and meta-regression findings cautiously, especially when few studies are available.
Identify the specific unresolved source of heterogeneity that matters scientifically or practically.
Design the new study so that it directly tests or clarifies that source of variation.
Avoid collecting another generic dataset if it would merely add another unexplained estimate.