Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

When Does a Lack of Comparable Studies Make Synthesis Impossible?

Studies do not have to be identical to be synthesized, but they must be comparable in ways that matter to the research question. When differences undermine a meaningful common inference, statistical pooling may be inappropriate even though systematic review remains possible.

575
When Studies Cannot Be Synthesized Guide 575 of 760
01 · The Question

How Different Can Studies Be Before Combining Them Stops Making Sense?

You find several studies related to your research question, but comparison quickly becomes awkward. Participants differ substantially. Researchers define the intervention or exposure differently. Outcomes are measured using different instruments, follow-up periods range from weeks to years, and some studies use entirely different designs.

There is clearly a literature. What is less clear is whether those studies provide evidence about sufficiently similar questions to be synthesized meaningfully.

This creates an important distinction. A lack of comparability can make a particular form of synthesis, especially meta-analysis, inappropriate or impossible. It does not necessarily make systematic review impossible.

02 · The Short Answer

Studies Need Meaningful Comparability, Not Perfect Similarity

In Brief

A lack of comparable studies prevents a particular synthesis when the studies differ so fundamentally in the questions they address, the constructs they measure, or the evidence they provide that combining their results would not support a coherent interpretation.

That threshold is not defined by a single heterogeneity statistic or by the presence of different methods. Studies that cannot be meta-analyzed may still be systematically reviewed and synthesized using other methods. The key question is what can legitimately be inferred across them.

03 · What You Need to Know

Comparability Depends on the Question You Are Trying to Answer

Two studies are not comparable simply because they share keywords, and they do not become incomparable merely because their methods differ. Comparability is relational: studies must be sufficiently aligned with the particular question and synthesis you intend to conduct.

For intervention research, researchers often examine populations, interventions, comparators, and outcomes. Other research questions may require attention to exposures, constructs, contexts, time points, study designs, or effect measures. The relevant dimensions depend on the substantive question.

Cochrane distinguishes clinical diversity, methodological diversity, and statistical heterogeneity. These forms of variation require judgment rather than a mechanical threshold. Its guidance explicitly recommends thoughtful consideration of whether numerical results are appropriate to combine before performing a meta-analysis.

Different but meaningfully comparable Studies vary in some characteristics but still provide interpretable evidence about a sufficiently coherent underlying question.
Fundamentally incomparable for the proposed synthesis Differences alter what is being studied or estimated so substantially that a common summary would have unclear or misleading meaning.

Different Populations Do Not Automatically Prevent Synthesis

Suppose studies evaluate the same intervention among adolescents, university students, and working adults. Those populations differ, but whether they can be synthesized depends on the research question.

If your question concerns the average effect across a broad population and there is a defensible reason to consider the effects related, synthesis may be appropriate. If your question specifically concerns adolescents, pooling adult populations into the estimate may answer a different question.

Population diversity can also be scientifically informative. Instead of asking whether the studies are “too different,” you may need to ask whether population differences plausibly modify the effect and whether the available evidence can investigate that possibility.

Different Measures Can Sometimes Be Harmonized, but Constructs Still Matter

Studies often measure the same construct using different instruments. In some circumstances, statistical methods such as standardized mean differences permit quantitative synthesis across different scales intended to measure the same underlying outcome.

But mathematical standardization cannot establish conceptual equivalence.

For example, scores from two validated measures of depressive symptoms may sometimes be treated as measures of a sufficiently similar construct for a particular synthesis. Combining examination performance, student satisfaction, course completion, and self-reported motivation simply because all four were described as “learning outcomes” would require a much stronger substantive justification.

Before converting numbers to a common metric, establish that the underlying outcomes are meaningfully related to the question being asked.

Different Interventions or Exposures May Represent Different Questions

Broad labels can conceal substantial differences. “Online learning,” “artificial intelligence use,” “feedback,” or “exercise intervention” may encompass interventions that operate through very different mechanisms.

If one study evaluates automated feedback on essays and another evaluates an AI tutoring system for mathematics, placing both under “AI in education” does not automatically make their effect estimates meaningfully combinable.

The problem is conceptual before it is statistical. A pooled estimate is useful only if you can explain what the resulting average represents.

Different Study Designs Require More Than Putting Results in the Same Table

Randomized trials, quasi-experiments, cross-sectional surveys, longitudinal observational studies, case-control studies, and qualitative research produce different kinds of evidence. They can certainly appear within a broad systematic review when the question warrants it, but they should not be combined indiscriminately.

Even among quantitative studies, effect estimates may represent different causal or associational quantities. Statistical conversion does not automatically make their inferential meaning equivalent.

Sometimes the appropriate response is to organize studies into conceptually defensible groups and synthesize each group separately. The fact that all studies appear in one review does not require one pooled answer.

Meta-Analysis May Be Impossible Even When Systematic Synthesis Is Not

Cochrane identifies several situations in which meta-analysis of effect estimates may not be possible or appropriate. These include limited evidence, incompletely reported outcomes or effect estimates, concerns about bias, and substantial clinical or methodological diversity. Statistical heterogeneity may also make a single average misleading in some circumstances.

None of these automatically requires abandoning the systematic review.

Other synthesis methods may be appropriate, depending on the available data and question. Cochrane's guidance on synthesis without meta-analysis emphasizes that these methods should be specified explicitly rather than hidden under vague labels such as “narrative synthesis.” Vote counting based only on whether individual results are statistically significant is specifically discouraged because it can produce misleading conclusions.

Cannot meta-analyze A defensible statistical pooling of effect estimates is unavailable or inappropriate for the particular comparison.
Cannot synthesize meaningfully The available studies are so disconnected from one another or from the intended question that even a structured cross-study conclusion would require claims the evidence cannot support.

Too Much Statistical Heterogeneity Is Not a Universal Stop Rule

Researchers sometimes conclude that a high I2 value means meta-analysis is prohibited. That is too mechanical.

Cochrane notes that statistical heterogeneity requires careful interpretation and that presenting an average combined effect can be misleading, particularly when study effects include both benefit and harm. Possible responses include checking data and effect measures, investigating heterogeneity where appropriate, presenting prediction intervals, modifying comparisons with justification, or using another synthesis method.

The broader issue is explored when deciding whether heterogeneous existing evidence justifies another primary study. Heterogeneity should provoke investigation, not merely trigger an arbitrary numerical cutoff.

Sometimes Narrowing the Synthesis Solves the Comparability Problem

Suppose 25 studies are relevant to a broad question, but only eight evaluate sufficiently similar interventions and outcomes for a particular meta-analysis. You do not necessarily need to choose between pooling all 25 and abandoning quantitative synthesis entirely.

A defensible review can organize studies into meaningful comparisons and synthesize only those groups for which a common inference makes sense. Other studies can remain part of the systematic review and contribute through another appropriate synthesis method.

This is one reason eligibility criteria and synthesis groups should be planned carefully. Decisions made after seeing study results can introduce considerable analytical flexibility, so departures from prespecified plans should be justified transparently.

Insufficient Comparability Can Reveal a Standardization Problem

When every study uses a different outcome definition, time point, intervention specification, or reporting format, the inability to synthesize may itself reveal something important about the field.

The next contribution might therefore be methodological rather than another generic study. Researchers could work toward common outcome definitions, stronger measurement, reporting standards, or more consistent intervention descriptions.

If new primary research is warranted, it should ideally reduce rather than reproduce the comparability problem.

Incomparability Can Also Reveal a Genuine Data Gap

Sometimes the existing studies simply do not answer the same question you need answered. Perhaps the relevant population appears in only one study, the critical outcome has rarely been measured, or existing interventions are materially different from the one now being considered.

At that point, the limitation is no longer merely “we cannot pool these studies.” The evidence required for your question may genuinely be missing.

This connects directly with deciding whether to collect new data when existing studies have not been synthesized properly. First determine whether synthesis can extract the needed information. If the necessary observations do not exist, new primary research becomes much easier to justify.

Do Not Force a Summary Merely Because Meta-Analysis Software Can Produce One

Statistical software can combine estimates whenever you supply numbers in compatible formats. That is a computational capability, not a scientific justification.

Before pooling, you should be able to state clearly what the summary effect means, which studies it represents, why those studies belong in the same comparison, and what variation remains relevant to interpretation.

Watch Out

Do not use statistical compatibility as a substitute for conceptual comparability. Two effect estimates can sometimes be converted into the same numerical format while still representing substantively different questions.

04 · A Practical Example

When One Literature Contains Several Different Questions

Hypothetical Example

Twenty studies of “AI-supported learning” that are not one intervention

Suppose a researcher identifies 20 studies examining AI-supported learning in higher education and initially plans a meta-analysis of academic performance.

Initial evidence pool The studies all mention AI-supported learning, but some evaluate intelligent tutoring systems, others automated writing feedback, and others general-purpose generative AI tools.
Closer comparison Outcomes include examination scores, assignment grades, course completion, self-reported learning, and satisfaction. Designs range from randomized experiments to cross-sectional surveys.
Synthesis decision The researcher concludes that one overall effect of “AI-supported learning” would have no sufficiently coherent interpretation. However, a subset of six controlled studies evaluates comparable tutoring interventions using objective achievement outcomes.
Revised synthesis Those six studies are considered for quantitative synthesis, while the broader evidence is organized and synthesized using methods appropriate to its different questions and data.

The literature was not “unsynthesizable” in an absolute sense. The original proposed synthesis was simply too broad. Once the research question and comparison were refined, some evidence could be combined and other evidence could still be reviewed without forcing everything into one pooled estimate.

05 · What Researchers Often Get Wrong

Common Mistakes When Studies Are Difficult to Compare

Misconception

“Different Measures Mean Meta-Analysis Is Impossible”

Not necessarily. Different instruments may sometimes measure a sufficiently similar construct and can, in appropriate circumstances, be represented using a common effect measure. The crucial question is conceptual comparability, not whether the original scales have identical names.

Misconception

“High Heterogeneity Means I Must Abandon the Meta-Analysis”

There is no universal I2 threshold that automatically determines whether pooling is scientifically defensible. Examine the effects, their direction, the study differences, uncertainty, and the purpose of the synthesis before deciding how to proceed.

Misconception

“If Meta-Analysis Is Impossible, the Systematic Review Has Failed”

No. Meta-analysis is one possible synthesis method within a systematic review. A review can remain valuable without statistical pooling if it uses transparent, appropriate methods to synthesize and present the available evidence.

Misconception

“I Can Solve Comparability by Converting Everything to an Effect Size”

Common numerical scales solve some statistical differences, not substantive ones. Standardizing a mathematics test score and a motivation scale does not make them measurements of the same outcome.

Misconception

“If I Cannot Pool the Studies, I Should Count How Many Are Significant”

Vote counting based on statistical significance is misleading because significance depends partly on sample size and precision. Cochrane explicitly identifies this approach as unacceptable. When meta-analysis is unavailable, use a synthesis method appropriate to the type of data rather than replacing effect estimation with a tally of p-values.

06 · What This Means for You

Decide What Can Be Combined Before Deciding That Nothing Can

When studies look incomparable, work from the research question outward. Identify the population, phenomenon, intervention or exposure, comparator, outcome, timing, and design features that matter to the inference you want to make.

You may discover that the full evidence base cannot support one synthesis but that meaningful subsets can. You may instead find that systematic synthesis without meta-analysis is appropriate. Only sometimes will the studies be so disconnected that the specific cross-study inference you wanted cannot be supported at all.

A simple decision framework

If studies differ but still address a coherent underlying question
Consider whether differences can be accommodated through appropriate synthesis methods, grouping, effect measures, or explicit modelling of variation.
If only some studies are meaningfully comparable
Form defensible synthesis groups rather than forcing the entire evidence base into one analysis.
If quantitative pooling would be misleading but studies still inform the same broader question
Use an appropriate synthesis method without meta-analysis and report it explicitly.
If studies measure fundamentally different constructs or answer materially different questions
Do not manufacture a common summary. Reframe the synthesis or acknowledge that the intended inference is unsupported.
If the evidence needed for the question simply does not exist
Consider targeted primary research designed to produce the missing comparable evidence.

Sometimes the inability to synthesize is not an analytical inconvenience. It is evidence about how fragmented the field has become. That finding can be quite useful, even if it does not produce the satisfying diamond at the bottom of a forest plot.

07 · A Quick Checklist

Before Declaring Studies Too Different to Synthesize

Before abandoning synthesis, check:
Define the exact cross-study question you want the synthesis to answer.
Compare the studies on the populations, interventions or exposures, comparators, outcomes, time points, and designs relevant to that question.
Distinguish conceptual incompatibility from differences that can legitimately be handled statistically or through structured grouping.
Determine whether meaningful subsets of studies can be synthesized separately.
Consider synthesis methods other than meta-analysis when quantitative pooling is inappropriate or unavailable.
Avoid vote counting based solely on whether individual studies report statistically significant results.
Explain transparently why studies were or were not combined and what inference remains possible.
If new research is needed, design it to improve comparability rather than reproducing the fragmentation identified in the review.
08 · Frequently Asked Questions

Questions About Study Comparability and Evidence Synthesis

Do studies have to use the same methods to be synthesized?

No. Methodological differences are common. What matters is whether the studies can contribute meaningfully to the particular synthesis and whether differences are handled appropriately rather than ignored.

Can studies using different outcome scales be meta-analyzed?

Sometimes. Different scales measuring sufficiently comparable constructs may be synthesized using appropriate effect measures. Statistical conversion does not justify combining outcomes that represent substantively different constructs.

What I² value means studies are too heterogeneous to combine?

There is no universal cutoff that settles the decision. Interpretation depends on the magnitude and direction of effects, the evidence for heterogeneity, study characteristics, and the question the pooled estimate is intended to answer.

Can I conduct a systematic review if I know meta-analysis will not be possible?

Yes. Systematic reviews do not require meta-analysis. Other synthesis methods can be used when appropriate, provided the methods are transparently specified and suited to the evidence.

Should I use narrative synthesis when studies cannot be pooled?

A non-meta-analytic synthesis may be appropriate, but simply labeling the analysis “narrative synthesis” is not enough. Specify how studies were grouped, what data were used, how findings were compared, and how conclusions were derived. The SWiM reporting guideline can be useful for reviews using synthesis methods without meta-analysis.

Does a lack of comparable studies justify another primary study?

It can, particularly when the evidence needed to answer an important question is genuinely missing. The new study should be designed to supply that evidence and, where possible, improve compatibility with other high-quality research.

Can I exclude inconvenient studies so the remaining studies are homogeneous?

Not simply because their results complicate the analysis. Eligibility and synthesis decisions should follow defensible criteria related to the research question and methodology. Post hoc changes should be transparently justified rather than driven by a desire for a cleaner pooled result.

09 · The Bottom Line

Do Not Confuse “Cannot Pool” With “Cannot Learn”

The Bottom Line

A lack of comparable studies makes a particular synthesis impossible or misleading when differences are so consequential that the studies cannot support a coherent common inference, but failure of meta-analysis does not automatically mean failure of systematic synthesis.

Refine the question, identify meaningful synthesis groups, and consider appropriate non-meta-analytic methods before concluding that the evidence cannot be combined. If the required comparable evidence genuinely does not exist, that limitation can provide a much stronger rationale for targeted primary research.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes