01 · The Question
How Can Two Systematic Reviews Look at the Same Evidence and Disagree?
You find one systematic review concluding that an intervention is beneficial. Another review published around the same period describes the evidence as uncertain. A third finds little convincing effect. If systematic reviews are supposed to synthesize the evidence systematically, how can all three exist?
The apparent contradiction often comes from assuming that reviews with similar titles are answering exactly the same question using exactly the same evidence and methods. Frequently, they are not. Differences can enter at almost every stage between defining the question and writing the conclusion.
03 · What You Need to Know
Where Systematic Reviews Begin to Diverge
The Reviews May Not Actually Ask the Same Question
Two titles can look nearly identical while the underlying review questions differ in consequential ways. One review may include only adults while another includes adolescents. One may define an intervention narrowly while another groups several related interventions together. Outcomes, comparators, follow-up periods, settings, and eligible study designs may also differ.
Before comparing conclusions, compare the questions. If the underlying questions differ, apparently conflicting answers may both be reasonable within their respective scopes.
Different answers to different questions
Reviews differ because their populations, interventions or exposures, comparators, outcomes, settings, timing, or eligible designs are not equivalent.
Different answers to substantially the same question
Reviews address closely comparable questions but diverge because of search coverage, study selection, appraisal, analysis, or interpretation.
Searches Conducted at Different Times See Different Evidence
A review searched in 2021 and another searched in 2025 do not necessarily have access to the same evidence base. New studies may have appeared, unpublished work may have become available, and corrections or retractions may have changed what should be considered.
Publication date alone can obscure this difference because reviews may take considerable time to move from the final search to publication. Compare the final search dates rather than merely asking which paper appeared most recently.
Even reviews conducted at similar times can retrieve different evidence if they search different databases, registers, languages, or supplementary sources. This is why determining whether each review searched comprehensively enough can help explain disagreement.
Eligibility Criteria Can Produce Different Collections of Studies
One review may include randomized trials only. Another may include randomized and non-randomized studies. One may accept several outcome measures while another requires a particular validated measure. Restrictions involving publication status, language, duration, intervention characteristics, or study setting can further change which evidence enters the synthesis.
Neither the broader nor narrower review is automatically superior. Eligibility should correspond to the review question and intended inference. What matters is whether the criteria are defensible and consistently applied.
When two reviews disagree, compare which primary studies actually entered each review and why .
Reviews Often Overlap Without Containing Identical Evidence
Multiple reviews on the same topic frequently include many of the same primary studies while also containing studies absent from the others. Cochrane identifies overlapping systematic reviews as a common methodological challenge in overviews of reviews and recommends mapping which primary studies occur in which reviews.
This overlap matters because two reviews described as independent summaries may actually be drawing heavily on the same underlying evidence. Conversely, a small number of non-overlapping studies may be responsible for their different results.
Watch Out
Do not count several overlapping systematic reviews as though they were several independent bodies of evidence. The same primary study may appear repeatedly across reviews, sometimes contributing substantial weight to each synthesis.
Reviews Can Group the Same Studies Differently
Even when two reviews identify the same primary studies, they may make different decisions about which studies belong in the same synthesis. One review may pool several intervention variants together, while another analyzes them separately. Outcomes may be grouped differently, follow-up periods may be combined or separated, and different effect measures may be chosen.
These decisions change what the pooled estimate represents. If one review calculates a broad average while another separates clinically or substantively different comparisons, their numerical results may not be directly comparable.
The key question is whether the studies grouped together form a meaningful synthesis , not simply whether both reviews contain a meta-analysis.
Different Statistical Models Can Produce Different Estimates
Meta-analysis involves analytical choices. Reviewers may use fixed-effect or random-effects models, different estimators of between-study variance, alternative effect measures, different approaches to zero events or missing statistics, and different assumptions about dependence among effect estimates.
These choices need not indicate manipulation. Several methods may be defensible. Their consequences become important when conclusions depend heavily on which reasonable analytical option is chosen.
A useful question is therefore whether the conclusion is robust across plausible analytical choices. If a minor methodological change moves the synthesis from a clear benefit to substantial uncertainty, that sensitivity is itself important information.
Reviews Can Treat Heterogeneity Differently
Suppose the same studies produce considerable variation in effect sizes. One review may emphasize the pooled average. Another may focus on the variation and conclude that no single effect adequately represents all settings.
Both could report the same underlying data while framing their conclusions differently. This is particularly important when high heterogeneity complicates interpretation of the pooled estimate .
Risk-of-Bias Judgments Can Differ
Critical appraisal involves structured judgment. Review teams may use different tools, versions of tools, domain definitions, or interpretations of incomplete reporting. They may consequently rate the same study differently.
More importantly, reviews can differ in what they do with those judgments. One may conduct sensitivity analyses excluding studies with serious concerns. Another may pool everything together and mention risk of bias only in the discussion.
AMSTAR 2 treats assessment of risk of bias and consideration of that bias when interpreting review results as critical methodological domains.
Publication Bias and Missing Evidence May Be Handled Differently
Reviews can differ in whether they search registers or unpublished sources, compare protocols with reports, investigate small-study effects, or assess possible selective reporting. One synthesis may therefore contain evidence that another never identified.
If the availability of studies or results depends on what they found, publication bias can distort the resulting meta-analysis . Reviews that handle this risk differently may reasonably reach different levels of confidence in otherwise similar pooled estimates.
The Numerical Results May Agree More Than the Authors' Conclusions Do
Sometimes the disagreement is primarily interpretive. Two reviews may report similar effect estimates and confidence intervals yet use noticeably different language in their abstracts and conclusions.
One team may describe a modest statistically significant effect as evidence of benefit. Another may emphasize that the effect is small, evidence certainty is limited, or practical importance remains uncertain. A third may regard the same evidence as insufficient for a strong recommendation.
When reviews appear to conflict, compare the actual effect estimates and uncertainty before comparing the prose. Abstract conclusions can make a modest interpretive difference look like a scientific standoff.
07 · A Quick Checklist
How to Compare Systematic Reviews That Reach Different Conclusions
When reviews disagree, compare:
Are they actually asking the same research question?
When was the final search conducted in each review?
Which databases, registers, and supplementary sources did each review search?
How do their eligibility criteria differ?
Which primary studies overlap, and which studies appear in only one review?
Did they group populations, interventions, outcomes, or follow-up periods differently?
Did they use different meta-analytic models, effect measures, or sensitivity analyses?
How did each review assess and incorporate risk of bias, heterogeneity, and missing evidence?
Do the numerical results genuinely conflict, or is the disagreement mainly in interpretation?
11 · Cite this Guide
How to Cite This Guide
This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.
Recommended (Field Guide)
APA
MLA
Chicago
Copy Citation