Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How Do You Know Whether the Included Studies Were Too Different to Combine?

Studies do not need to be identical to be combined in a meta-analysis. The real question is whether their differences still allow the pooled estimate to represent a scientifically meaningful quantity.

425
Were the Studies Too Different to Combine? Guide 425 of 899
01 · The Question

When Do Differences Between Studies Make Pooling Misleading?

No two studies are identical. Participants differ, interventions vary, researchers measure outcomes differently, and studies are conducted in different settings. Some variation is therefore expected whenever research is synthesized.

The difficult question is not whether the studies differ. They almost certainly do. The question is whether they differ in ways that make a single pooled estimate scientifically difficult to interpret or potentially misleading.

02 · The Short Answer

Studies Are Too Different When Their Average Stops Answering a Coherent Question

In Brief

Studies are too different to combine when their differences in populations, interventions or exposures, comparators, outcomes, methods, or underlying effects make the proposed pooled estimate scientifically difficult to interpret as a meaningful summary.

There is no single statistical cutoff that decides this. Clinical and methodological comparability should be considered before pooling, while statistical heterogeneity helps reveal how much observed effects vary once comparable studies are analyzed together.

03 · What You Need to Know

Similarity Is a Scientific Judgment Before It Is a Statistical Test

Studies Do Not Need to Be Identical

Meta-analysis exists partly because separate studies differ. Combining sufficiently comparable studies can improve precision, summarize an evidence base, and reveal patterns that individual studies cannot establish alone. Cochrane identifies improved precision and the ability to address questions across studies among the potential advantages of meta-analysis.

Demanding identical studies would therefore defeat much of the purpose of evidence synthesis. Some diversity is expected and can even make a review more informative by showing how findings behave across contexts.

Ask Whether the Studies Estimate Effects Worth Averaging

The central question is conceptual: what does the pooled number mean?

If studies investigate reasonably related versions of an intervention in populations for whom an average effect is meaningful, pooling may be defensible despite differences. If one study evaluates a brief intervention in children, another a year-long program in working adults, and another a substantially different intervention in hospitalized patients, a common label may not be enough to make their average informative.

Cochrane advises considering meta-analysis only when a group of studies is sufficiently homogeneous in participants, interventions, and outcomes to provide a meaningful summary.

Clinical Diversity Comes First

Clinical diversity refers broadly to differences in characteristics such as participants, interventions, comparators, and outcomes. In non-clinical fields, the same principle applies even if "clinical" is an awkward label: studies may differ in learners, technologies, organizational settings, exposures, implementation, or outcome definitions.

Ask whether these differences could plausibly change the effect being estimated. Age may matter for one intervention but not another. A modest difference in dose may be trivial in one context and decisive in another.

This means that comparability cannot be judged from labels alone. Two studies may both claim to examine "online learning" while evaluating interventions so different that averaging them obscures more than it reveals.

Methodological Diversity Can Change Observed Effects

Studies may also differ in design, risk of bias, measurement instruments, follow-up periods, analytical methods, or definitions of outcomes. Cochrane describes this as methodological diversity and notes that such differences may produce variation in observed intervention effects.

For example, combining randomized and seriously confounded observational evidence into one undifferentiated estimate may require considerable justification. Likewise, combining outcomes measured with instruments that do not represent sufficiently similar constructs can create a pooled quantity that is statistically computable but scientifically vague.

Clinical diversity Differences in participants, interventions or exposures, comparators, outcomes, settings, or other substantive characteristics.
Methodological diversity Differences in study design, risk of bias, outcome measurement, analytical procedures, or other methodological characteristics.
Statistical heterogeneity Variation in observed effect estimates beyond what would be expected from sampling error alone.

Do Not Wait for I² Before Thinking About Comparability

A common mistake is to pool first and then use I² to decide whether pooling was appropriate. That reverses the logic.

Researchers should first decide whether the studies address a sufficiently coherent comparison for a pooled result to have meaning. Statistical heterogeneity is then examined as part of understanding variation among those studies.

A low I² cannot make fundamentally incompatible interventions or outcomes conceptually equivalent. Conversely, a high I² does not automatically prove that pooling was inappropriate. It tells you that effect estimates vary more than expected from sampling error, and that variation requires interpretation.

I² Is Not a Universal Traffic Light

I² describes the proportion of variability in effect estimates attributable to heterogeneity rather than sampling error. It is useful, but its interpretation depends on the magnitude and direction of effects and the strength of evidence for heterogeneity.

Cochrane explicitly warns against using simple thresholds diagnostically and notes substantial uncertainty in I² and related measures when few studies are available.

So an I² of 60% does not mechanically mean "do not pool," just as an I² of 10% does not certify that pooling is sensible.

The Direction of Effects Can Matter More Than a Single Heterogeneity Statistic

Suppose all studies suggest benefit but differ in how large that benefit is. An average effect may still have a useful interpretation, depending on the question.

Now suppose some studies indicate substantial benefit while others indicate meaningful harm. An average close to zero could conceal a much more consequential pattern. Cochrane specifically cautions that when results vary considerably, particularly in the direction of effect, reporting an average intervention effect may be misleading.

Watch Out

A pooled effect near zero can mean "little effect everywhere," but it can also result from averaging substantial benefit in some circumstances with substantial harm in others. Those interpretations are not equivalent.

A Random-Effects Model Does Not Solve Incompatibility

Random-effects meta-analysis allows the underlying effects estimated by studies to differ and summarizes the centre of a distribution of effects. It is therefore useful in many heterogeneous evidence bases.

But switching from a fixed-effect to a random-effects model does not answer the prior scientific question of whether the studies should be combined at all. Cochrane explicitly states that random-effects analysis is not a substitute for investigating heterogeneity.

If studies estimate effects that do not form a meaningful family of questions, a more flexible statistical model cannot manufacture conceptual coherence.

Prediction Intervals Can Reveal What an Average Conceals

When random-effects meta-analysis is appropriate, a confidence interval around the pooled mean describes uncertainty about that mean. It does not directly show how widely the underlying effects vary across settings.

A prediction interval can sometimes provide a more intuitive representation of between-study variation by indicating a range within which the effect of a future similar study might be expected to lie, subject to the model and available evidence. Cochrane recommends considering prediction intervals as a way to present the extent of between-study variation.

This can be particularly informative when the average effect looks favorable but plausible effects across settings vary substantially.

Sometimes the Correct Decision Is Not to Pool

A systematic review does not need to contain a meta-analysis. If the studies are too diverse for an average effect to have a useful interpretation, reviewers can synthesize findings using other structured methods.

Cochrane explicitly lists not performing a meta-analysis as an option when variation is considerable, particularly when effects differ in direction.

Not producing a diamond at the bottom of a forest plot is not a methodological failure. Sometimes restraint is the more informative analysis. Even statisticians must occasionally resist a shiny diamond.

04 · A Practical Example

When the Same Intervention Label Hides Very Different Studies

Hypothetical Example

Pooling Studies of "AI-Assisted Learning"

Suppose a systematic review identifies eight studies described as evaluating AI-assisted learning and considers combining their effects on academic performance.

Inspect the participants Four studies involve undergraduate students, two involve secondary-school students, and two involve postgraduate professional training.
Inspect the interventions Some studies evaluate adaptive tutoring used throughout a semester, while others evaluate one-time generative AI feedback or automated practice questions.
Inspect the outcomes Several measure final examination performance, others use researcher-developed quizzes, and one reports course grades affected by several components unrelated to the intervention.
Inspect the designs The studies also differ substantially in allocation methods and risk of bias.
Ask what the pooled effect would mean An average labelled "the effect of AI-assisted learning" may imply a single intervention category and outcome that the underlying studies do not actually share.

The studies might still support useful synthesis. Reviewers could potentially define more coherent groups, examine intervention types separately, or use a structured synthesis without calculating one overall effect. The important decision comes from the scientific meaning of the comparisons, not from whether software is capable of calculating a pooled estimate.

05 · What Researchers Often Get Wrong

Common Mistakes When Deciding Whether Studies Can Be Combined

Misconception

Any Studies on the Same Topic Can Be Pooled

Sharing a broad topic does not establish that studies estimate sufficiently related effects. The proposed pooled estimate needs a defensible scientific interpretation.

Misconception

A Low I² Proves the Studies Belong Together

No. I² describes observed statistical heterogeneity; it does not determine whether populations, interventions, outcomes, and methods are conceptually comparable. That judgment must be made using substantive and methodological knowledge.

Misconception

A High I² Automatically Means You Cannot Pool

Also no. High heterogeneity is an important signal requiring investigation and cautious interpretation, but its implications depend on the magnitude and direction of effects, possible explanations, and the purpose of the synthesis. The separate question is how high heterogeneity should change your trust in the pooled estimate.

Misconception

Random Effects Fixes Heterogeneity

A random-effects model incorporates an assumption about variation in underlying effects. It does not remove heterogeneity or explain why effects differ, and it cannot make scientifically incompatible studies comparable.

Misconception

More Studies Make Pooling Easier to Justify

More studies provide more information, but they may also reveal greater diversity. A larger meta-analysis is not automatically stronger if the additional studies make the pooled question less coherent or introduce serious methodological limitations.

06 · What This Means for You

Decide What You Are Averaging Before You Average It

When reading a meta-analysis, begin with the substantive comparison rather than the heterogeneity statistic. Identify the studies contributing to the pooled estimate and ask whether their differences still leave you with a meaningful average.

A simple decision framework

If populations, interventions, comparators, and outcomes represent a coherent scientific comparison
Pooling may be reasonable, subject to methodological quality and statistical heterogeneity.
If important substantive differences could plausibly change the effect
Examine whether separate analyses or prespecified subgroup comparisons provide a more interpretable synthesis.
If methodological differences are substantial
Consider whether different biases or measurement approaches make a common pooled estimate difficult to interpret.
If effects vary greatly or point in opposite directions
Investigate the variation and consider whether reporting one average would conceal more than it clarifies.
If no scientifically meaningful pooled quantity can be defined
Do not assume meta-analysis is required. A structured synthesis without statistical pooling may be more appropriate.

If pooling remains defensible but substantial statistical heterogeneity is present, the appraisal is not finished. You then need to decide how that heterogeneity changes the interpretation of the pooled result.

07 · A Quick Checklist

Before Accepting a Pooled Estimate, Check Whether the Studies Belong Together

Before interpreting the meta-analysis, check:
Do the studies address a sufficiently coherent research question for an average effect to have a clear meaning?
Are the populations sufficiently comparable, or are important population differences likely to modify the effect?
Are the interventions or exposures genuinely comparable rather than merely sharing a broad label?
Do the outcomes measure sufficiently similar constructs at comparable time points?
Could differences in study design, measurement, or risk of bias explain important differences in results?
Do individual effects point in broadly compatible directions, or would the average conceal benefit in some contexts and harm in others?
Were heterogeneity statistics interpreted in context rather than using a rigid I² cutoff?
If random effects were used, did the authors still investigate and discuss heterogeneity rather than treating the model as a solution?
Would separate analyses or synthesis without pooling communicate the evidence more meaningfully?
08 · Frequently Asked Questions

Questions About Combining Different Studies in Meta-Analysis

Do studies have to use exactly the same intervention to be combined?

No. Some intervention variation can be compatible with a meaningful synthesis. The key question is whether the variants represent effects that make scientific sense to summarize together for the review's intended inference.

What level of I² means studies should not be combined?

There is no universal I² threshold that determines whether pooling is appropriate. Cochrane warns that simple thresholds can be misleading because interpretation depends on the magnitude and direction of effects, evidence for heterogeneity, and uncertainty in I² itself.

Can studies with different populations be combined?

Sometimes. It depends on whether an average across those populations answers a meaningful question and whether population differences plausibly modify the effect. Broad synthesis can be legitimate when the scope and interpretation reflect that breadth.

Can different outcome scales be combined?

Sometimes, particularly when different instruments measure the same underlying construct and an appropriate effect measure such as a standardized mean difference is justified. Statistical convertibility alone is not enough; the outcomes still need sufficient conceptual similarity.

Does a random-effects model mean heterogeneous studies can always be pooled?

No. Random-effects models allow underlying effects to vary, but Cochrane emphasizes that they are not substitutes for investigating heterogeneity. The studies still need to form a scientifically meaningful synthesis.

What should reviewers do if studies are too different to combine?

They can present results without calculating one overall pooled effect, organize studies into more coherent comparisons, or use appropriate structured synthesis methods. Cochrane explicitly recognizes not performing a meta-analysis as a legitimate response to problematic heterogeneity.

Does heterogeneity mean the studies are poor quality?

Not necessarily. Effects can legitimately differ across populations, settings, or intervention variants even when studies are rigorous. Methodological differences and bias can also generate heterogeneity, so its likely sources need to be investigated rather than assumed.

09 · The Bottom Line

The Question Is Not Whether Studies Differ, but Whether Their Average Means Anything

The Bottom Line

Studies are too different to combine when their substantive or methodological differences make a pooled estimate scientifically difficult to interpret as a meaningful answer to the review question.

Judge clinical and methodological comparability before relying on I² or another heterogeneity statistic. Statistical models can describe or accommodate some variation, but they cannot turn fundamentally incompatible evidence into a coherent scientific question.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes