01 · The Question
Are you synthesizing the evidence, or synthesizing what authors said about it?
Literature reviews often inherit their conclusions from the papers they summarize. Study A “demonstrated” that an intervention works. Study B “confirmed” that one variable influences another. A review repeats the conclusion. Later papers cite the review. Eventually, the interpretation can become more familiar than the evidence underneath it.
But the results and the authors' interpretation of those results are not the same thing.
An observational association can be described causally. A statistically significant but trivial difference can be presented as important. A wide confidence interval can become a confident conclusion. Findings from one population can be generalized much further than the sample permits. Meta-research refers to some forms of distorted or overly favorable interpretation as spin, including conclusions unsupported by the results, unsupported causal attribution, and inappropriate extrapolation.
The question is therefore fundamental: if you temporarily removed the authors' discussion and conclusion, what would the design, data, estimates, and uncertainty actually allow you to say?
03 · What You Need to Know
How do you determine what a study actually establishes?
Separate four layers that papers often compress together
A useful reading habit is to distinguish what was done, what was observed, what can reasonably be inferred, and what the authors ultimately conclude.
| Layer |
Question to ask |
| Design and measurement |
What was actually studied, compared, manipulated, observed, and measured? |
| Results |
What estimates, patterns, differences, associations, or qualitative findings were actually obtained? |
| Inference |
What conclusions are justified given the design, bias, precision, and alternative explanations? |
| Interpretation |
How do the authors explain, generalize, frame, or assign importance to those findings? |
Problems arise when the fourth layer is treated as though it were simply another result.
The study design sets boundaries around the claim
Suppose a cross-sectional survey finds that students who report greater generative AI use also report lower confidence in their independent writing ability.
The observed association is a result. The claim that AI use reduces writing confidence is a causal interpretation.
That interpretation requires assumptions beyond the observed association. Temporal order may be unclear. Students with lower confidence might use AI more frequently. Other variables could influence both AI use and confidence. Measurement or selection could also matter.
The appropriate question is not whether causal language sounds plausible. It is whether the design provides the evidence necessary to support it.
The abstract conclusion deserves checking against the results
Abstracts are especially influential because many readers encounter the conclusion before examining the full methods and results. Research on scientific “spin” has documented discordance between results and interpretation, inappropriate extrapolation, and causal language unsupported by study design. Experimental work has also shown that overstated abstract conclusions can affect readers' interpretation of research findings.
When a study matters to your synthesis, do not extract only the conclusion sentence. Compare it with the primary results and the design that produced them.
Watch Out
The discussion section is not where results acquire additional statistical power. If the methods and results support a cautious conclusion, confident prose several pages later does not upgrade the evidence.
Statistical significance does not tell you what the evidence means
A p-value addresses a statistical question under a specified model. It does not tell you whether the observed effect is large enough to matter, whether the study is unbiased, whether the result generalizes, or whether the causal explanation is correct.
Cochrane recommends focusing interpretation on effect estimates and confidence intervals rather than dividing findings into “statistically significant” and “non-significant” categories. It notes that a small p-value can accompany an effect too small to be important, while wide intervals can leave substantial uncertainty despite an apparently favorable point estimate.
So replace “the study found an effect” with a more informative question: what was the estimated effect, how precise was it, and what substantively important possibilities remain compatible with the evidence?
Association does not become causation through repetition
If ten papers repeat the sentence “X influences Y” while all ultimately rely on observational associations, the tenth repetition does not supply the missing causal identification.
Repeated interpretation can make a proposition feel established even when the underlying evidence has not changed. This is why consequential claims should sometimes be traced back to their original evidence.
Once there, evaluate what the study actually measured and what alternative explanations remain.
Authors may generalize beyond the studied population
A study conducted among 200 first-year students at one university can provide evidence about those participants under the studied conditions. Whether its findings generalize to university students internationally is another inference.
The authors may reasonably discuss broader implications, but your synthesis should distinguish evidence directly observed in the study from extrapolation to populations, settings, or conditions not examined.
This does not make generalization illegitimate. It makes generalization something that requires justification.
A mechanism can be plausible without being demonstrated
Discussion sections often propose explanations for observed findings. These can be intellectually valuable and generate important research questions.
But if the mechanism itself was not measured or experimentally tested, describe it accordingly.
Observed finding
The measured pattern, estimate, difference, association, or qualitative evidence produced by the study.
Proposed explanation
A theoretically or empirically plausible account of why the finding occurred that may require additional evidence to establish.
A discussion can suggest why an effect occurred. That suggestion should not quietly become “the study showed that the effect occurred because...” in your literature review.
A study can establish one result while leaving the broader theory unresolved
Theories often generate several predictions. Evidence consistent with one prediction may support the theory without uniquely establishing it, particularly when competing theories predict the same observation.
Ask whether the result discriminates among explanations or merely remains compatible with one of them.
This distinction becomes important when a literature repeatedly describes findings as “supporting” a theory. Support can range from weak compatibility to a demanding test that rules out credible alternatives. Those are not equivalent achievements.
Secondary sources can inherit and strengthen the original interpretation
A review may summarize an observational study as showing that X predicts Y. A later paper may describe the review as evidence that X affects Y. Another may state that X causes Y.
No new empirical evidence is necessarily added during that progression.
Tracing important claims to their source can reveal whether later certainty reflects genuinely stronger subsequent research or simply rhetorical drift. When many publications ultimately rely on the same underlying evidence, you should also ask whether the conclusion depends heavily on one study, dataset, research group, or method.
Author conclusions can also be too pessimistic
Overinterpretation is not always optimistic.
Authors can describe a non-significant result as showing “no effect” even when the confidence interval remains compatible with meaningful benefit or harm. Cochrane specifically identifies confusion between lack of evidence of an effect and evidence of no effect as a common interpretive error.
This is why absence of evidence must be distinguished from evidence of absence.
Interpretation should reflect certainty as well as the point estimate
A point estimate is not the complete finding. Its uncertainty and the credibility of the evidence matter.
Cochrane recommends calibrating conclusions according to both estimated effects and certainty of evidence. When certainty is very low, for example, the appropriate conclusion may be that the evidence is very uncertain about the effect rather than asserting the point estimate as though it were established.
The same principle applies outside formal GRADE assessments. Strong language should require strong evidential justification.
Your literature review should synthesize evidence, not vote on author conclusions
Suppose six papers conclude that an intervention is effective and two conclude that it is not. Counting those conclusions produces an eight-paper opinion poll.
Instead, examine the studies beneath them. Perhaps the six favorable papers are small uncontrolled studies while the two cautious papers are large randomized trials. Perhaps the opposite is true. Perhaps all eight estimate broadly similar effects but use different verbal thresholds for calling them important.
This is where explaining why some evidence deserves more weight than other evidence becomes more useful than counting conclusions.
04 · A Practical Example
How a plausible conclusion can outrun the study that produced it
Hypothetical Example
“Generative AI weakens students' critical thinking”
Suppose a cross-sectional study surveys 3,000 university students. Students report how frequently they use generative AI for coursework and complete a short self-report measure intended to capture aspects of critical-thinking disposition.
Higher reported AI use is modestly associated with lower scores on the measure after adjustment for several demographic variables. The authors conclude that reliance on generative AI weakens students' critical-thinking abilities.
The study has produced evidence of an adjusted association between self-reported AI use and scores on a particular measure at one time point.
It has not by itself established that AI use caused the difference, that actual critical-thinking performance declined, or that reducing AI use would improve critical thinking. Those propositions may be plausible and worthy of testing, but they require additional evidence.
What was measured?
Self-reported AI use and a particular measure of critical-thinking disposition.
What was observed?
A modest adjusted cross-sectional association.
What remains unresolved?
Temporal order, residual confounding, measurement validity, and causal direction.
What can be concluded?
The study supports an association under the observed conditions and motivates stronger causal investigation.
What cannot yet be concluded?
That AI use itself caused a decline in students' critical-thinking ability.
07 · A Quick Checklist
Does your claim stop where the evidence stops?
Before repeating an author's conclusion, check:
I understand what was actually measured, compared, manipulated, or observed.
I have examined the relevant effect estimates and uncertainty rather than relying only on significance labels or conclusion sentences.
The causal strength of my wording matches what the study design can support.
I distinguish directly observed findings from mechanisms or explanations proposed in the discussion.
I have not generalized beyond the studied population, outcome, setting, or time period without additional justification.
For repeated claims, I know whether independent evidence has actually accumulated or whether later papers mainly repeat earlier interpretations.
My interpretation reflects relevant risk of bias, precision, directness, and certainty rather than the point estimate alone.
If I removed the authors' discussion section, I could still explain why my conclusion follows from the methods and results.