03 · What You Need to Know
What to examine when critically appraising a research paper
Start with the research question
Before criticizing the statistics or sample size, establish what the researchers were actually trying to find out.
Ask:
- Is the research question or objective clear?
- Is the question specific enough to evaluate?
- Does the study address a meaningful scientific or practical problem?
- Are the population, phenomenon, exposure, intervention, comparison, or outcome defined where relevant?
- Do the hypotheses match the stated question?
A beautifully executed study can still have limited usefulness if its question is poorly defined. Conversely, you cannot evaluate whether a design is appropriate until you know what question the design is supposed to answer.
Ask whether the study design can answer that question
This is one of the most consequential appraisal questions.
Different designs support different kinds of inference. A cross-sectional observational study can identify associations at a particular time but generally cannot, by itself, establish the temporal sequence needed for many causal claims. A randomized controlled trial can provide stronger evidence about intervention effects when randomization, adherence, outcome measurement, missing data, analysis, and other aspects are handled appropriately.
A qualitative study may be appropriate for understanding experiences, meanings, processes, and perspectives even though it is not designed to estimate a population-level causal effect.
Do not rank designs mechanically without considering the research question. Ask instead:
Is this design appropriate for the claim the researchers want to make?
Examine how participants, cases, or data were selected
A study's conclusions depend partly on who or what generated the evidence.
Look for:
- the target population;
- the actual study population;
- inclusion and exclusion criteria;
- recruitment or sampling procedures;
- response or participation patterns;
- important differences between participants and nonparticipants;
- loss to follow-up where relevant; and
- whether the sample represents the population to which the authors generalize.
Selection problems can distort both internal validity and applicability. A sample can be perfectly real and accurately measured while still providing weak evidence for a broader population if the selection process systematically favors particular participants.
Do not ask only whether the sample is “large enough”
Sample size matters, but bigger is not automatically better.
A large biased sample can produce a very precise estimate of the wrong quantity. A smaller well-designed study may provide more informative evidence for a focused question.
Ask why the sample size was appropriate. Depending on the design, consider whether the researchers reported an a priori sample-size or power calculation, whether enough outcome events occurred for the planned analysis, and whether attrition reduced the effective sample substantially.
Then examine the uncertainty around the actual estimates rather than treating the original sample-size calculation as proof that the study was adequately informative.
Ask whether the important variables were measured well
A study cannot rescue a badly measured variable through sophisticated analysis.
For major exposures, interventions, outcomes, predictors, and constructs, ask:
- How was the variable defined?
- How was it measured?
- Was the measurement method validated for this purpose or population where appropriate?
- Was measurement consistent across groups?
- Could participants or assessors have known information that influenced measurement?
- Were self-reported measures appropriate?
- Could misclassification or measurement error materially affect the results?
A familiar questionnaire name or laboratory measure should not end the appraisal. Check whether it measures the construct the authors claim it measures in the context in which they used it.
Look for bias as a process, not as a vague criticism
“The study may be biased” is not a useful appraisal unless you can explain how.
Bias is a systematic deviation from the effect or quantity the study is trying to estimate. Different study designs create different opportunities for bias.
For randomized trials, Cochrane's RoB 2 framework considers domains including the randomization process, deviations from intended interventions, missing outcome data, outcome measurement, and selection of the reported result.
For non-randomized intervention studies, ROBINS-I considers bias due to confounding, participant selection, intervention classification, deviations from intended interventions, missing data, outcome measurement, and selection of the reported result.
When you identify a possible bias, ask three questions:
- What mechanism could create the bias?
- In what direction might it affect the result?
- Is the problem large enough to change the conclusion?
Distinguish bias from random error
Bias
A systematic problem that can move an estimate away from the quantity the study is trying to estimate.
Random error
Sampling or measurement variation that creates uncertainty around an estimate and can occur even in an otherwise well-designed study.
A larger sample can often reduce sampling uncertainty. It does not automatically eliminate systematic bias.
This distinction explains why an enormous dataset does not guarantee a trustworthy causal estimate. Precision and validity are related but different properties.
For observational studies, examine confounding carefully
Suppose an observational study reports that people who use a particular technology have better academic performance. The users and nonusers may also differ in prior achievement, socioeconomic circumstances, motivation, age, access to resources, or other characteristics related to the outcome.
Those differences can create or distort an association.
Ask:
- What plausible confounders exist?
- Which were measured?
- How accurately were they measured?
- How did the analysis address them?
- Were important confounders omitted?
- Could residual confounding remain after adjustment?
Statistical adjustment is not a magic conversion from observational association to causal proof. Adjustment works only with the variables measured and the assumptions built into the analysis.
For randomized trials, examine the randomization process rather than accepting the label
The words “randomized controlled trial” are not themselves evidence that randomization worked properly.
Look for how the allocation sequence was generated and concealed, whether important baseline imbalances raise concerns, whether participants remained analyzed in appropriate groups, and whether deviations from intended interventions could have affected the result.
Blinding may also matter, particularly when knowledge of allocation can influence behavior, co-interventions, outcome assessment, or reporting. But blinding is not feasible or equally important in every intervention.
Use design-specific appraisal rather than applying “Was it blinded?” as a universal quality test.
Examine missing data and loss to follow-up
Missing observations are not automatically harmless.
The important questions are why data are missing, whether missingness differs among groups, whether it is related to outcomes or exposures, how much data are missing, and how the analysis handled the problem.
A small percentage of missing data can sometimes be consequential if the missingness is highly informative. A larger percentage may be less damaging under different assumptions and analytical handling.
Do not evaluate missing data by percentage alone.
Check whether the statistical analysis matches the data and design
You do not need to reproduce every calculation to ask useful questions about an analysis.
Start with:
- Does the statistical method match the outcome and study design?
- Are important assumptions acknowledged or checked?
- Are repeated or clustered observations handled appropriately?
- Are confounders or covariates selected sensibly?
- Were important subgroup analyses planned or exploratory?
- Was missing data handled appropriately?
- Are estimates accompanied by measures of uncertainty?
- Do the reported analyses correspond to the research questions?
If a statistical method is central to the paper and you do not understand it, that is a reason to investigate the method rather than automatically trust or reject the result.
Look beyond the p-value
A p-value does not tell you whether a result is large, important, unbiased, clinically meaningful, practically useful, or likely to replicate.
Examine the estimated effect and its uncertainty.
Depending on the study, this might include:
- mean differences;
- risk differences;
- risk ratios;
- odds ratios;
- hazard ratios;
- correlations;
- regression coefficients;
- standardized effect sizes; or
- other discipline-appropriate estimates.
Confidence intervals can help you see which effect sizes remain reasonably compatible with the data under the model and assumptions used. A result can be statistically significant yet too small to matter practically, or statistically non-significant while still leaving substantial uncertainty about effects that would matter.
Check the actual numbers, tables, and figures
Do not evaluate a paper solely through the authors' prose.
Inspect the tables and figures yourself. Check whether:
- sample sizes are consistent;
- denominators change unexpectedly;
- effect estimates match the textual description;
- confidence intervals support the interpretation;
- axes or scales could exaggerate visual differences;
- important null findings are visible but underemphasized;
- subgroup results are based on very small numbers; and
- footnotes reveal qualifications not obvious from the main text.
If you are still developing an efficient reading routine, focus first on the high-information parts of the paper before trying to inspect every detail at once.
Separate statistical significance from practical importance
Large studies can detect very small differences. Whether those differences matter is a separate question.
Ask:
If this estimate is approximately correct, would it change a scientific explanation, clinical decision, policy, engineering choice, educational intervention, or other real decision?
The answer depends on context. There is no universal effect-size threshold that makes a result important.
Check whether the authors are making causal claims
Words matter.
Statements such as “X caused Y,” “X improved Y,” or “X reduced Y” imply something stronger than “X was associated with Y.”
When causal language appears, ask whether the study design and assumptions can support it.
Watch Out
Do not allow causal wording to outrun the design. An association can remain after statistical adjustment without proving that changing the exposure would change the outcome. Evaluate temporality, confounding, selection, measurement, and the design's causal assumptions before accepting causal interpretation.
Compare the conclusion with the results, not with the title
Titles and abstracts are highly compressed. Conclusions can also become more confident than the evidence warrants.
After reading the results, ask:
- Does the conclusion answer the original research question?
- Does it acknowledge uncertainty?
- Does it distinguish association from causation?
- Does it generalize beyond the studied population?
- Does it emphasize a secondary or subgroup result more than the primary result?
- Does it imply practical importance that was never demonstrated?
- Does it acknowledge important contradictory findings?
A useful appraisal often consists of rewriting the conclusion in language proportional to the evidence.
Look for selective reporting
Researchers make many analytical decisions: which outcomes to emphasize, which time points to report, which covariates to include, which subgroups to show, and which statistical models to present.
If these choices are made after examining the results and only favorable analyses are highlighted, the published picture can be misleading.
When available, compare the paper with its preregistration, trial registration, protocol, or statistical analysis plan. Check whether prespecified primary outcomes and analyses match what appears in the publication.
Cochrane includes selection of the reported result as an explicit risk-of-bias domain for randomized trials.
Read the limitations, but conduct your own appraisal
Authors' limitation sections are useful because they show which problems the researchers themselves recognize.
But critical appraisal does not mean copying that paragraph into your notes.
Look for limitations the authors may not emphasize, and distinguish limitations that merely exist from limitations that materially weaken the conclusion.
A paper does not become invalid because it has limitations. Every study has constraints. The question is what those constraints mean for the claim you want to take from it.
Check conflicts of interest and funding
Read the funding statement and conflict-of-interest disclosures.
A declared financial relationship does not automatically invalidate a study, and absence of commercial funding does not guarantee unbiased research.
Instead, ask whether relationships could plausibly influence study design, comparison choice, data access, analysis, interpretation, reporting, or publication. Then consider the safeguards and transparency provided.
The International Committee of Medical Journal Editors emphasizes disclosure of relationships and activities that readers could perceive as influencing the work.
Check whether the reporting is complete enough to evaluate
Poor reporting and poor research are not identical, but incomplete reporting can make trustworthy appraisal impossible.
Reporting guidelines can help you determine what information should normally be available for particular study designs. The EQUATOR Network maintains a searchable library of health-research reporting guidelines, including CONSORT for randomized trials, STROBE for observational research, PRISMA for systematic reviews, and many design- or topic-specific extensions.
These guidelines primarily improve reporting rather than serving as simple quality scores. A paper can mention every checklist item and still contain methodological problems. Conversely, missing reporting may leave you uncertain about what was actually done.
Consider whether the findings apply beyond the study
Internal validity asks whether the study provides a credible answer within the conditions examined. External validity or applicability asks whether that answer is useful elsewhere.
Consider:
- How similar is your target population to the study population?
- Would the intervention, exposure, or setting operate similarly?
- Are relevant demographic, institutional, cultural, geographic, or temporal differences important?
- Were participants unusually selected?
- Would implementation conditions differ?
A study can be internally strong but only narrowly applicable. That is not necessarily a flaw. It means you should state the boundaries of the evidence correctly.
Place the paper in the wider evidence base
One paper rarely settles an important scientific question.
Ask whether the finding is consistent with previous research, whether credible studies disagree, whether the paper replicates earlier evidence, and whether systematic reviews or other syntheses provide a broader picture.
Do not reject a surprising result simply because it disagrees with previous research. Equally, do not overturn an established evidence base because one new paper has a dramatic conclusion.
Use individual studies as pieces of evidence whose importance depends partly on their design, validity, precision, directness, and relationship to other evidence.
Do not confuse journal prestige with study validity
A journal's reputation, indexing status, Impact Factor, or quartile can tell you something about the publication venue. None substitutes for evaluating the study itself.
Peer review can improve manuscripts and filter some problems, but publication does not certify every analysis or conclusion as correct.
Read the paper, not merely the journal name.