01 · The Question
How far can a conclusion legitimately travel beyond the data?
A study reports a difference between two groups. The conclusion says an intervention works. A survey identifies an association. The authors conclude that one factor influences another. Interview participants describe a recurring difficulty. The paper concludes that an institutional practice should change.
Each conclusion may be defensible. But none is simply a repetition of the underlying result.
Research conclusions are inferential claims. They connect findings to judgments about explanation, causation, importance, generalizability, theory, practice, or future research. Critical reading requires you to see that connection rather than allowing the conclusion to replace the evidence that supposedly supports it.
03 · What You Need to Know
How results become conclusions
Start by writing down the result without the conclusion
Before evaluating a conclusion, establish what the study actually found before the authors interpreted it.
For a quantitative study, this might be an estimated difference between groups, an association between variables, a change over time, a model coefficient, or an estimated effect with uncertainty. For qualitative research, it may be a pattern, theme, process, account, or interpretation developed through the analytical approach.
Then state the authors’ conclusion separately.
Putting the two claims next to each other often reveals an inferential step that is difficult to notice when reading the paper continuously.
What the data show
The reported observations, estimates, patterns, comparisons, or analytical findings produced within the study.
What the authors conclude
The claim the researchers believe those findings justify about explanation, causation, importance, applicability, theory, practice, or another broader issue.
Every conclusion contains an inferential bridge
Consider a study finding that students who report greater use of a learning platform also obtain higher examination scores.
The empirical result might be an association between platform use and examination performance. Several conclusions could then be proposed:
- greater platform use is associated with higher examination scores;
- using the platform improves academic performance;
- the platform improves performance because it increases engagement;
- universities should require students to use the platform.
These statements do not merely become longer versions of the same result. Each introduces additional claims.
The first stays close to the observed association. The second introduces causation. The third adds a causal mechanism. The fourth moves into a recommendation that may require evidence about benefits, harms, feasibility, costs, alternatives, and context.
A useful appraisal question is therefore: What had to be assumed or established to move from the result to this conclusion?
The study design limits what can reasonably be concluded
Different designs support different kinds of inference. A randomized experiment, prospective cohort, cross-sectional survey, case-control study, qualitative interview study, and systematic review do not provide interchangeable evidence.
For example, an association observed in an observational study may be compatible with several explanations, including causal effects, confounding, selection processes, measurement problems, or other biases. The strength of a causal conclusion depends on the design and the assumptions required for causal inference.
This is why you should pay close attention when authors shift from association to causation. A grammatical change from “was associated with” to “led to” can represent a substantial change in the scientific claim.
Statistical results do not establish substantive importance by themselves
A numerical difference can be statistically detectable while remaining too small to matter for the practical question being studied. Conversely, an estimate that does not cross a conventional significance threshold can still be compatible with effects that would matter in practice.
Therefore, a conclusion such as “the intervention produced an important improvement” contains at least two judgments: that the analysis supports a difference and that the magnitude of that difference is important.
The latter requires substantive criteria. It cannot be supplied by the p-value alone. When the paper equates the two, examine whether it is treating statistical significance as practical importance.
Uncertainty should survive the trip into the conclusion
Researchers sometimes report uncertainty carefully in the Results section and then write a conclusion that sounds considerably more certain.
Suppose an effect estimate is compatible with anything from a trivial benefit to a substantial one. A conclusion claiming that the intervention “produces substantial improvement” suppresses information contained in that uncertainty.
Likewise, failure to obtain statistically significant evidence does not ordinarily establish that there is no effect. The distinction becomes important when authors move from an inconclusive result to a categorical statement of equivalence or ineffectiveness without an appropriate design and analysis. That is the difference between having no evidence of an effect and establishing no effect.
Population claims can extend beyond the sample
A study may collect data from a narrowly defined group and then discuss a much broader population.
Whether that generalization is defensible depends on sampling, setting, eligibility criteria, context, mechanisms, and other considerations related to external validity or transferability. Statistical inference from a sample does not automatically establish that the same result applies across populations, institutions, countries, age groups, or circumstances that were not adequately represented.
Whenever the population named in the conclusion is broader than the population studied, notice the expansion.
A finding about a measured variable may become a claim about a larger construct
Measurement creates another possible gap between data and conclusion.
Suppose a study uses a short self-report scale as its measure of “engagement.” The analysis directly concerns scores produced by that measure. A broad conclusion about student engagement assumes that the instrument adequately represents the construct for the population and context under study.
That assumption may be well supported. It still deserves recognition.
The same issue appears when researchers move from test performance to “learning,” from publication counts to “research productivity,” from self-reported intention to actual behavior, or from one operational definition to a broader theoretical construct.
Mechanisms require evidence beyond the outcome itself
An intervention can produce an outcome without establishing why it produced that outcome.
If students receiving a new teaching method perform better, that result alone does not demonstrate that increased motivation caused the improvement. Motivation would need to be measured and the proposed mechanism evaluated using an appropriate design and analysis before such a claim becomes substantially stronger.
Possible explanations are perfectly legitimate in a Discussion section when identified as possibilities. Problems arise when plausible explanations quietly become established mechanisms.
Recommendations travel farther than descriptive conclusions
Evidence that something occurred is not necessarily sufficient evidence that a particular action should follow.
A recommendation may require consideration of benefits, harms, feasibility, cost, equity, alternatives, stakeholder values, implementation conditions, and the broader evidence base. A single study can contribute to that judgment without settling it.
STROBE guidance for observational research explicitly recommends cautious interpretation that considers the study objectives, limitations, multiplicity of analyses, findings from related studies, and other relevant evidence. It also emphasizes viewing an individual study as a contribution to a wider literature rather than automatically treating it as a stand-alone basis for action.
The conclusion should not contain a new empirical finding
The Conclusion section should synthesize what follows from findings already reported. It should not suddenly introduce a result that was absent from the Results section.
If you encounter a striking claim only at the end of the paper, trace it backward. Find the result that supports it. If you cannot, you may be dealing with a conclusion that extends beyond the reported results.
07 · A Quick Checklist
Does the conclusion stay within the evidence?
Before accepting an author conclusion, check:
Write down the specific finding or findings that support the conclusion.
Identify whether the conclusion is descriptive, associational, causal, mechanistic, predictive, practical, or prescriptive.
Check whether the study design supports that type of inference.
Compare the certainty of the conclusion with the uncertainty in the reported results.
Check whether the conclusion generalizes beyond the studied sample, population, setting, outcome, or time period.
Separate demonstrated mechanisms from explanations that remain hypothetical.
Verify that major claims in the conclusion correspond to findings reported earlier in the paper.
Ask what additional assumptions or evidence are required to move from the finding to the final claim.