03 · What You Need to Know
Similarity Is a Scientific Judgment Before It Is a Statistical Test
Studies Do Not Need to Be Identical
Meta-analysis exists partly because separate studies differ. Combining sufficiently comparable studies can improve precision, summarize an evidence base, and reveal patterns that individual studies cannot establish alone. Cochrane identifies improved precision and the ability to address questions across studies among the potential advantages of meta-analysis.
Demanding identical studies would therefore defeat much of the purpose of evidence synthesis. Some diversity is expected and can even make a review more informative by showing how findings behave across contexts.
Ask Whether the Studies Estimate Effects Worth Averaging
The central question is conceptual: what does the pooled number mean?
If studies investigate reasonably related versions of an intervention in populations for whom an average effect is meaningful, pooling may be defensible despite differences. If one study evaluates a brief intervention in children, another a year-long program in working adults, and another a substantially different intervention in hospitalized patients, a common label may not be enough to make their average informative.
Cochrane advises considering meta-analysis only when a group of studies is sufficiently homogeneous in participants, interventions, and outcomes to provide a meaningful summary.
Clinical Diversity Comes First
Clinical diversity refers broadly to differences in characteristics such as participants, interventions, comparators, and outcomes. In non-clinical fields, the same principle applies even if "clinical" is an awkward label: studies may differ in learners, technologies, organizational settings, exposures, implementation, or outcome definitions.
Ask whether these differences could plausibly change the effect being estimated. Age may matter for one intervention but not another. A modest difference in dose may be trivial in one context and decisive in another.
This means that comparability cannot be judged from labels alone. Two studies may both claim to examine "online learning" while evaluating interventions so different that averaging them obscures more than it reveals.
Methodological Diversity Can Change Observed Effects
Studies may also differ in design, risk of bias, measurement instruments, follow-up periods, analytical methods, or definitions of outcomes. Cochrane describes this as methodological diversity and notes that such differences may produce variation in observed intervention effects.
For example, combining randomized and seriously confounded observational evidence into one undifferentiated estimate may require considerable justification. Likewise, combining outcomes measured with instruments that do not represent sufficiently similar constructs can create a pooled quantity that is statistically computable but scientifically vague.
Clinical diversity
Differences in participants, interventions or exposures, comparators, outcomes, settings, or other substantive characteristics.
Methodological diversity
Differences in study design, risk of bias, outcome measurement, analytical procedures, or other methodological characteristics.
Statistical heterogeneity
Variation in observed effect estimates beyond what would be expected from sampling error alone.
Do Not Wait for I² Before Thinking About Comparability
A common mistake is to pool first and then use I² to decide whether pooling was appropriate. That reverses the logic.
Researchers should first decide whether the studies address a sufficiently coherent comparison for a pooled result to have meaning. Statistical heterogeneity is then examined as part of understanding variation among those studies.
A low I² cannot make fundamentally incompatible interventions or outcomes conceptually equivalent. Conversely, a high I² does not automatically prove that pooling was inappropriate. It tells you that effect estimates vary more than expected from sampling error, and that variation requires interpretation.
I² Is Not a Universal Traffic Light
I² describes the proportion of variability in effect estimates attributable to heterogeneity rather than sampling error. It is useful, but its interpretation depends on the magnitude and direction of effects and the strength of evidence for heterogeneity.
Cochrane explicitly warns against using simple thresholds diagnostically and notes substantial uncertainty in I² and related measures when few studies are available.
So an I² of 60% does not mechanically mean "do not pool," just as an I² of 10% does not certify that pooling is sensible.
The Direction of Effects Can Matter More Than a Single Heterogeneity Statistic
Suppose all studies suggest benefit but differ in how large that benefit is. An average effect may still have a useful interpretation, depending on the question.
Now suppose some studies indicate substantial benefit while others indicate meaningful harm. An average close to zero could conceal a much more consequential pattern. Cochrane specifically cautions that when results vary considerably, particularly in the direction of effect, reporting an average intervention effect may be misleading.
Watch Out
A pooled effect near zero can mean "little effect everywhere," but it can also result from averaging substantial benefit in some circumstances with substantial harm in others. Those interpretations are not equivalent.
A Random-Effects Model Does Not Solve Incompatibility
Random-effects meta-analysis allows the underlying effects estimated by studies to differ and summarizes the centre of a distribution of effects. It is therefore useful in many heterogeneous evidence bases.
But switching from a fixed-effect to a random-effects model does not answer the prior scientific question of whether the studies should be combined at all. Cochrane explicitly states that random-effects analysis is not a substitute for investigating heterogeneity.
If studies estimate effects that do not form a meaningful family of questions, a more flexible statistical model cannot manufacture conceptual coherence.
Prediction Intervals Can Reveal What an Average Conceals
When random-effects meta-analysis is appropriate, a confidence interval around the pooled mean describes uncertainty about that mean. It does not directly show how widely the underlying effects vary across settings.
A prediction interval can sometimes provide a more intuitive representation of between-study variation by indicating a range within which the effect of a future similar study might be expected to lie, subject to the model and available evidence. Cochrane recommends considering prediction intervals as a way to present the extent of between-study variation.
This can be particularly informative when the average effect looks favorable but plausible effects across settings vary substantially.
Sometimes the Correct Decision Is Not to Pool
A systematic review does not need to contain a meta-analysis. If the studies are too diverse for an average effect to have a useful interpretation, reviewers can synthesize findings using other structured methods.
Cochrane explicitly lists not performing a meta-analysis as an option when variation is considerable, particularly when effects differ in direction.
Not producing a diamond at the bottom of a forest plot is not a methodological failure. Sometimes restraint is the more informative analysis. Even statisticians must occasionally resist a shiny diamond.
07 · A Quick Checklist
Before Accepting a Pooled Estimate, Check Whether the Studies Belong Together
Before interpreting the meta-analysis, check:
Do the studies address a sufficiently coherent research question for an average effect to have a clear meaning?
Are the populations sufficiently comparable, or are important population differences likely to modify the effect?
Are the interventions or exposures genuinely comparable rather than merely sharing a broad label?
Do the outcomes measure sufficiently similar constructs at comparable time points?
Could differences in study design, measurement, or risk of bias explain important differences in results?
Do individual effects point in broadly compatible directions, or would the average conceal benefit in some contexts and harm in others?
Were heterogeneity statistics interpreted in context rather than using a rigid I² cutoff?
If random effects were used, did the authors still investigate and discuss heterogeneity rather than treating the model as a solution?
Would separate analyses or synthesis without pooling communicate the evidence more meaningfully?