03 · What You Need to Know
How do you determine whether two studies are genuinely comparable?
Translate each study into the question it actually answered
Do not begin with the title or the authors' concluding sentence. Reconstruct the empirical question from the methods.
Who was studied? What exposure or intervention occurred? Compared with what? Which outcome was measured? When was it measured? Under what conditions? What design produced the estimate?
For intervention reviews, Cochrane uses the PICO framework of population, intervention, comparator, and outcomes to define the question and emphasizes planning how different populations, interventions, outcomes, and study designs will be grouped for synthesis.
Population
Who or what does the result describe?
Exposure or intervention
What condition, experience, treatment, or factor is being examined?
Comparator
What is it being compared against?
Outcome
What exactly was measured?
Time and design
When was the outcome measured, and what kind of inference can the design support?
Once those elements are visible, many supposed contradictions become less mysterious.
The same population label can hide very different participants
Two studies of “university students” may involve populations with little practical similarity.
One may recruit first-year nursing students at a highly selective university. Another may recruit adult distance learners across several institutions. A third may study postgraduate engineering students with extensive prior experience in the technology being evaluated.
Age, prior knowledge, socioeconomic conditions, language, educational level, institutional context, baseline risk, motivation, and many other characteristics can modify how a phenomenon manifests or how an intervention works.
Cochrane notes that variation in participant characteristics is a form of clinical diversity and may generate heterogeneity when those characteristics affect the intervention effect.
A shared intervention label does not guarantee a shared intervention
Broad labels are especially dangerous in literature synthesis.
“Flipped classroom,” “blended learning,” “gamification,” “AI-assisted learning,” “peer feedback,” and “simulation” can describe families of interventions rather than standardized treatments.
Compare what participants actually experienced:
- What components were provided?
- How frequently?
- For how long?
- With what instructions?
- With what level of instructor involvement?
- Under what implementation conditions?
Two interventions sharing a noun phrase may differ enough that their effects need not be identical.
The comparator can quietly change the research question
“The intervention improved performance” is incomplete unless you know compared with what.
An AI tutoring system compared with no additional support addresses a different contrast from the same system compared with intensive human tutoring. An online course compared with no instruction is not the same question as an online course compared with a well-designed face-to-face course.
Different comparators can therefore produce different effect estimates without contradiction.
Intervention versus no intervention
Estimates what changes when the intervention is added relative to its absence.
Intervention versus active alternative
Estimates whether the intervention differs from another substantive approach.
Studies can use the same word for different outcomes
Construct labels deserve particular suspicion.
Consider “engagement.” One study measures login frequency. Another uses a self-report engagement scale. Another codes observable classroom behavior. Another measures persistence to course completion.
Those outcomes may be related, but they are not interchangeable.
The same problem occurs with achievement, learning, satisfaction, motivation, critical thinking, AI literacy, well-being, research productivity, and countless other constructs.
JBI appraisal guidance emphasizes whether outcomes are measured validly and reliably because differences in operationalization can alter what the study actually observes.
Immediate and long-term outcomes can legitimately differ
Time is part of the question.
An intervention may improve immediate test performance while having little effect on retention six months later. A behavioral change may appear during supervised implementation and disappear after support is removed. Adverse effects may emerge only after extended exposure.
Studies conducted at different follow-up times can therefore produce different estimates without one invalidating the other.
Ask what each time point represents rather than compressing all measurements into “the effect.”
Different research designs may estimate different things
Consider two studies examining social media use and depression. A cross-sectional study estimates an association between current social media use and current depressive symptoms. A longitudinal study asks whether earlier use predicts later symptoms. A randomized intervention reducing social media exposure estimates the effect of an assigned behavioral change under particular conditions.
These studies inhabit the same topic area but do not provide interchangeable answers.
Design also changes vulnerability to bias. JBI's current appraisal frameworks examine design-specific concerns such as confounding, temporal precedence, exposure classification, outcome measurement, participant retention, and statistical validity.
Adjusted and unadjusted estimates may answer different statistical questions
Even within similar observational designs, two analyses may not estimate the same relationship.
One paper may report the crude association between an exposure and outcome. Another adjusts for age, prior achievement, socioeconomic status, and baseline motivation. A third adjusts for a variable that may actually lie on the causal pathway.
Differences among these estimates can arise because the models encode different assumptions and statistical targets.
Before treating adjusted coefficients as competing estimates, examine which variables were included and why. JBI explicitly includes identification and handling of confounding as a core concern in appraising observational studies.
Different scales can make similar effects look different
Studies may express results using raw mean differences, standardized mean differences, odds ratios, risk ratios, correlations, regression coefficients, or other measures.
A numerical value of 0.30 does not have the same interpretation across all of these metrics. Even studies using the same effect measure may define the outcome direction differently.
Before declaring numerical disagreement, make sure the effect estimates are expressed on comparable scales and in comparable directions.
Authors' conclusions can disagree even when their results do not
This is one of the quieter sources of apparent conflict.
Two studies can report similar effect estimates but use very different language. One discussion calls the effect meaningful. Another calls it modest. One abstract emphasizes statistical significance. Another emphasizes the small magnitude.
If you compare only authors' prose, you may manufacture a disagreement that disappears when you compare the underlying results.
This is why you should eventually distinguish what the evidence establishes from what authors claim about it.
Statistical significance can manufacture apparent contradiction
Suppose Study A estimates an effect of 0.20 with a narrow confidence interval and reports p =.04. Study B estimates 0.18 with a wider confidence interval and reports p =.12.
It would be incorrect to conclude automatically that Study A found an effect while Study B found no effect and therefore the studies contradict one another. Their point estimates are extremely similar.
The difference lies largely in precision.
Watch Out
“Significant in one study but not significant in another” is not itself evidence that the effects differ. If the question is whether effects differ between groups or studies, that difference needs to be evaluated directly rather than inferred from separate significance tests.
Deciding what can be combined is a substantive judgment
Meta-analysis does not eliminate the need to decide whether studies estimate sufficiently comparable quantities.
Cochrane states that meta-analysis should be considered only when studies are sufficiently homogeneous in participants, interventions, and outcomes to provide a meaningful summary.
Its guidance also emphasizes planning how different populations, interventions, outcomes, and study designs will be grouped for synthesis.
Statistical software will happily calculate an average from numbers you give it. Whether that average answers a coherent research question remains your problem. Software has many talents; methodological embarrassment is not yet one of them.
Only after comparability is established should you explain genuine disagreement
Once you determine that studies address sufficiently similar questions and still produce materially different results, you have reached the next problem: genuine disagreement.
Then you can investigate sampling variation, bias, effect modification, implementation, analytical choices, and other possible explanations for why important studies disagree.
The order matters. Otherwise you may spend considerable analytical energy explaining a contradiction that was never there.