03 · What You Need to Know
Comparability Depends on the Question You Are Trying to Answer
Two studies are not comparable simply because they share keywords, and they do not become incomparable merely because their methods differ. Comparability is relational: studies must be sufficiently aligned with the particular question and synthesis you intend to conduct.
For intervention research, researchers often examine populations, interventions, comparators, and outcomes. Other research questions may require attention to exposures, constructs, contexts, time points, study designs, or effect measures. The relevant dimensions depend on the substantive question.
Cochrane distinguishes clinical diversity, methodological diversity, and statistical heterogeneity. These forms of variation require judgment rather than a mechanical threshold. Its guidance explicitly recommends thoughtful consideration of whether numerical results are appropriate to combine before performing a meta-analysis.
Different but meaningfully comparable
Studies vary in some characteristics but still provide interpretable evidence about a sufficiently coherent underlying question.
Fundamentally incomparable for the proposed synthesis
Differences alter what is being studied or estimated so substantially that a common summary would have unclear or misleading meaning.
Different Populations Do Not Automatically Prevent Synthesis
Suppose studies evaluate the same intervention among adolescents, university students, and working adults. Those populations differ, but whether they can be synthesized depends on the research question.
If your question concerns the average effect across a broad population and there is a defensible reason to consider the effects related, synthesis may be appropriate. If your question specifically concerns adolescents, pooling adult populations into the estimate may answer a different question.
Population diversity can also be scientifically informative. Instead of asking whether the studies are “too different,” you may need to ask whether population differences plausibly modify the effect and whether the available evidence can investigate that possibility.
Different Measures Can Sometimes Be Harmonized, but Constructs Still Matter
Studies often measure the same construct using different instruments. In some circumstances, statistical methods such as standardized mean differences permit quantitative synthesis across different scales intended to measure the same underlying outcome.
But mathematical standardization cannot establish conceptual equivalence.
For example, scores from two validated measures of depressive symptoms may sometimes be treated as measures of a sufficiently similar construct for a particular synthesis. Combining examination performance, student satisfaction, course completion, and self-reported motivation simply because all four were described as “learning outcomes” would require a much stronger substantive justification.
Before converting numbers to a common metric, establish that the underlying outcomes are meaningfully related to the question being asked.
Different Interventions or Exposures May Represent Different Questions
Broad labels can conceal substantial differences. “Online learning,” “artificial intelligence use,” “feedback,” or “exercise intervention” may encompass interventions that operate through very different mechanisms.
If one study evaluates automated feedback on essays and another evaluates an AI tutoring system for mathematics, placing both under “AI in education” does not automatically make their effect estimates meaningfully combinable.
The problem is conceptual before it is statistical. A pooled estimate is useful only if you can explain what the resulting average represents.
Different Study Designs Require More Than Putting Results in the Same Table
Randomized trials, quasi-experiments, cross-sectional surveys, longitudinal observational studies, case-control studies, and qualitative research produce different kinds of evidence. They can certainly appear within a broad systematic review when the question warrants it, but they should not be combined indiscriminately.
Even among quantitative studies, effect estimates may represent different causal or associational quantities. Statistical conversion does not automatically make their inferential meaning equivalent.
Sometimes the appropriate response is to organize studies into conceptually defensible groups and synthesize each group separately. The fact that all studies appear in one review does not require one pooled answer.
Meta-Analysis May Be Impossible Even When Systematic Synthesis Is Not
Cochrane identifies several situations in which meta-analysis of effect estimates may not be possible or appropriate. These include limited evidence, incompletely reported outcomes or effect estimates, concerns about bias, and substantial clinical or methodological diversity. Statistical heterogeneity may also make a single average misleading in some circumstances.
None of these automatically requires abandoning the systematic review.
Other synthesis methods may be appropriate, depending on the available data and question. Cochrane's guidance on synthesis without meta-analysis emphasizes that these methods should be specified explicitly rather than hidden under vague labels such as “narrative synthesis.” Vote counting based only on whether individual results are statistically significant is specifically discouraged because it can produce misleading conclusions.
Cannot meta-analyze
A defensible statistical pooling of effect estimates is unavailable or inappropriate for the particular comparison.
Cannot synthesize meaningfully
The available studies are so disconnected from one another or from the intended question that even a structured cross-study conclusion would require claims the evidence cannot support.
Too Much Statistical Heterogeneity Is Not a Universal Stop Rule
Researchers sometimes conclude that a high I2 value means meta-analysis is prohibited. That is too mechanical.
Cochrane notes that statistical heterogeneity requires careful interpretation and that presenting an average combined effect can be misleading, particularly when study effects include both benefit and harm. Possible responses include checking data and effect measures, investigating heterogeneity where appropriate, presenting prediction intervals, modifying comparisons with justification, or using another synthesis method.
The broader issue is explored when deciding whether heterogeneous existing evidence justifies another primary study. Heterogeneity should provoke investigation, not merely trigger an arbitrary numerical cutoff.
Sometimes Narrowing the Synthesis Solves the Comparability Problem
Suppose 25 studies are relevant to a broad question, but only eight evaluate sufficiently similar interventions and outcomes for a particular meta-analysis. You do not necessarily need to choose between pooling all 25 and abandoning quantitative synthesis entirely.
A defensible review can organize studies into meaningful comparisons and synthesize only those groups for which a common inference makes sense. Other studies can remain part of the systematic review and contribute through another appropriate synthesis method.
This is one reason eligibility criteria and synthesis groups should be planned carefully. Decisions made after seeing study results can introduce considerable analytical flexibility, so departures from prespecified plans should be justified transparently.
Insufficient Comparability Can Reveal a Standardization Problem
When every study uses a different outcome definition, time point, intervention specification, or reporting format, the inability to synthesize may itself reveal something important about the field.
The next contribution might therefore be methodological rather than another generic study. Researchers could work toward common outcome definitions, stronger measurement, reporting standards, or more consistent intervention descriptions.
If new primary research is warranted, it should ideally reduce rather than reproduce the comparability problem.
Incomparability Can Also Reveal a Genuine Data Gap
Sometimes the existing studies simply do not answer the same question you need answered. Perhaps the relevant population appears in only one study, the critical outcome has rarely been measured, or existing interventions are materially different from the one now being considered.
At that point, the limitation is no longer merely “we cannot pool these studies.” The evidence required for your question may genuinely be missing.
This connects directly with deciding whether to collect new data when existing studies have not been synthesized properly. First determine whether synthesis can extract the needed information. If the necessary observations do not exist, new primary research becomes much easier to justify.
Do Not Force a Summary Merely Because Meta-Analysis Software Can Produce One
Statistical software can combine estimates whenever you supply numbers in compatible formats. That is a computational capability, not a scientific justification.
Before pooling, you should be able to state clearly what the summary effect means, which studies it represents, why those studies belong in the same comparison, and what variation remains relevant to interpretation.
Watch Out
Do not use statistical compatibility as a substitute for conceptual comparability. Two effect estimates can sometimes be converted into the same numerical format while still representing substantively different questions.