Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How to Critically Evaluate a Research Paper: What Should You Actually Look For?

Critical appraisal is more than finding limitations. Learn how to judge whether a study asked a worthwhile question, used appropriate methods, produced credible results, and made conclusions the evidence can actually support.

08
How to Critically Evaluate a Research Paper Guide 8 of 13
01 · The Question

How do you decide whether a research paper is actually convincing?

A paper can be published in a respected journal, contain sophisticated statistics, and still provide weaker evidence than its conclusion suggests. Another study may have obvious limitations yet provide useful evidence for a carefully defined question.

Critical evaluation is therefore not about finding enough faults to declare a paper “good” or “bad.” It is about asking whether the research question, design, conduct, analysis, results, and interpretation fit together well enough to support the claims being made.

Different study designs require different appraisal questions. The appropriate concerns for a randomized trial are not identical to those for a cohort study, diagnostic-accuracy study, qualitative study, or systematic review. This is why established appraisal systems provide design-specific tools rather than one universal checklist. CASP, for example, provides separate checklists for multiple study types, while the Joanna Briggs Institute provides critical-appraisal tools tailored to different research designs.

The practical goal is to determine three things: Can you trust the study's findings, what exactly do those findings mean, and how far can you reasonably apply them?

02 · The Short Answer

Evaluate the chain from question to conclusion

In Brief

Ask whether the study posed a clear question, chose a design capable of answering it, selected participants or data appropriately, measured the important variables well, minimized important sources of bias and confounding, used suitable analyses, reported the results transparently, and drew conclusions proportional to the evidence.

Then consider uncertainty, limitations, applicability, conflicts of interest, and whether the findings fit with the wider evidence.

Do not reduce this process to a score based on how many checklist boxes a paper passes. Critical appraisal requires judgment about which weaknesses could materially change the result or its interpretation. Cochrane's risk-of-bias framework similarly focuses on specific mechanisms through which study design or conduct can bias findings rather than treating every reporting imperfection as equally consequential.

03 · What You Need to Know

What to examine when critically appraising a research paper

Start with the research question

Before criticizing the statistics or sample size, establish what the researchers were actually trying to find out.

Ask:

  • Is the research question or objective clear?
  • Is the question specific enough to evaluate?
  • Does the study address a meaningful scientific or practical problem?
  • Are the population, phenomenon, exposure, intervention, comparison, or outcome defined where relevant?
  • Do the hypotheses match the stated question?

A beautifully executed study can still have limited usefulness if its question is poorly defined. Conversely, you cannot evaluate whether a design is appropriate until you know what question the design is supposed to answer.

Ask whether the study design can answer that question

This is one of the most consequential appraisal questions.

Different designs support different kinds of inference. A cross-sectional observational study can identify associations at a particular time but generally cannot, by itself, establish the temporal sequence needed for many causal claims. A randomized controlled trial can provide stronger evidence about intervention effects when randomization, adherence, outcome measurement, missing data, analysis, and other aspects are handled appropriately.

A qualitative study may be appropriate for understanding experiences, meanings, processes, and perspectives even though it is not designed to estimate a population-level causal effect.

Do not rank designs mechanically without considering the research question. Ask instead:

Is this design appropriate for the claim the researchers want to make?

Examine how participants, cases, or data were selected

A study's conclusions depend partly on who or what generated the evidence.

Look for:

  • the target population;
  • the actual study population;
  • inclusion and exclusion criteria;
  • recruitment or sampling procedures;
  • response or participation patterns;
  • important differences between participants and nonparticipants;
  • loss to follow-up where relevant; and
  • whether the sample represents the population to which the authors generalize.

Selection problems can distort both internal validity and applicability. A sample can be perfectly real and accurately measured while still providing weak evidence for a broader population if the selection process systematically favors particular participants.

Do not ask only whether the sample is “large enough”

Sample size matters, but bigger is not automatically better.

A large biased sample can produce a very precise estimate of the wrong quantity. A smaller well-designed study may provide more informative evidence for a focused question.

Ask why the sample size was appropriate. Depending on the design, consider whether the researchers reported an a priori sample-size or power calculation, whether enough outcome events occurred for the planned analysis, and whether attrition reduced the effective sample substantially.

Then examine the uncertainty around the actual estimates rather than treating the original sample-size calculation as proof that the study was adequately informative.

Ask whether the important variables were measured well

A study cannot rescue a badly measured variable through sophisticated analysis.

For major exposures, interventions, outcomes, predictors, and constructs, ask:

  • How was the variable defined?
  • How was it measured?
  • Was the measurement method validated for this purpose or population where appropriate?
  • Was measurement consistent across groups?
  • Could participants or assessors have known information that influenced measurement?
  • Were self-reported measures appropriate?
  • Could misclassification or measurement error materially affect the results?

A familiar questionnaire name or laboratory measure should not end the appraisal. Check whether it measures the construct the authors claim it measures in the context in which they used it.

Look for bias as a process, not as a vague criticism

“The study may be biased” is not a useful appraisal unless you can explain how.

Bias is a systematic deviation from the effect or quantity the study is trying to estimate. Different study designs create different opportunities for bias.

For randomized trials, Cochrane's RoB 2 framework considers domains including the randomization process, deviations from intended interventions, missing outcome data, outcome measurement, and selection of the reported result.

For non-randomized intervention studies, ROBINS-I considers bias due to confounding, participant selection, intervention classification, deviations from intended interventions, missing data, outcome measurement, and selection of the reported result.

When you identify a possible bias, ask three questions:

  1. What mechanism could create the bias?
  2. In what direction might it affect the result?
  3. Is the problem large enough to change the conclusion?

Distinguish bias from random error

Bias A systematic problem that can move an estimate away from the quantity the study is trying to estimate.
Random error Sampling or measurement variation that creates uncertainty around an estimate and can occur even in an otherwise well-designed study.

A larger sample can often reduce sampling uncertainty. It does not automatically eliminate systematic bias.

This distinction explains why an enormous dataset does not guarantee a trustworthy causal estimate. Precision and validity are related but different properties.

For observational studies, examine confounding carefully

Suppose an observational study reports that people who use a particular technology have better academic performance. The users and nonusers may also differ in prior achievement, socioeconomic circumstances, motivation, age, access to resources, or other characteristics related to the outcome.

Those differences can create or distort an association.

Ask:

  • What plausible confounders exist?
  • Which were measured?
  • How accurately were they measured?
  • How did the analysis address them?
  • Were important confounders omitted?
  • Could residual confounding remain after adjustment?

Statistical adjustment is not a magic conversion from observational association to causal proof. Adjustment works only with the variables measured and the assumptions built into the analysis.

For randomized trials, examine the randomization process rather than accepting the label

The words “randomized controlled trial” are not themselves evidence that randomization worked properly.

Look for how the allocation sequence was generated and concealed, whether important baseline imbalances raise concerns, whether participants remained analyzed in appropriate groups, and whether deviations from intended interventions could have affected the result.

Blinding may also matter, particularly when knowledge of allocation can influence behavior, co-interventions, outcome assessment, or reporting. But blinding is not feasible or equally important in every intervention.

Use design-specific appraisal rather than applying “Was it blinded?” as a universal quality test.

Examine missing data and loss to follow-up

Missing observations are not automatically harmless.

The important questions are why data are missing, whether missingness differs among groups, whether it is related to outcomes or exposures, how much data are missing, and how the analysis handled the problem.

A small percentage of missing data can sometimes be consequential if the missingness is highly informative. A larger percentage may be less damaging under different assumptions and analytical handling.

Do not evaluate missing data by percentage alone.

Check whether the statistical analysis matches the data and design

You do not need to reproduce every calculation to ask useful questions about an analysis.

Start with:

  • Does the statistical method match the outcome and study design?
  • Are important assumptions acknowledged or checked?
  • Are repeated or clustered observations handled appropriately?
  • Are confounders or covariates selected sensibly?
  • Were important subgroup analyses planned or exploratory?
  • Was missing data handled appropriately?
  • Are estimates accompanied by measures of uncertainty?
  • Do the reported analyses correspond to the research questions?

If a statistical method is central to the paper and you do not understand it, that is a reason to investigate the method rather than automatically trust or reject the result.

Look beyond the p-value

A p-value does not tell you whether a result is large, important, unbiased, clinically meaningful, practically useful, or likely to replicate.

Examine the estimated effect and its uncertainty.

Depending on the study, this might include:

  • mean differences;
  • risk differences;
  • risk ratios;
  • odds ratios;
  • hazard ratios;
  • correlations;
  • regression coefficients;
  • standardized effect sizes; or
  • other discipline-appropriate estimates.

Confidence intervals can help you see which effect sizes remain reasonably compatible with the data under the model and assumptions used. A result can be statistically significant yet too small to matter practically, or statistically non-significant while still leaving substantial uncertainty about effects that would matter.

Check the actual numbers, tables, and figures

Do not evaluate a paper solely through the authors' prose.

Inspect the tables and figures yourself. Check whether:

  • sample sizes are consistent;
  • denominators change unexpectedly;
  • effect estimates match the textual description;
  • confidence intervals support the interpretation;
  • axes or scales could exaggerate visual differences;
  • important null findings are visible but underemphasized;
  • subgroup results are based on very small numbers; and
  • footnotes reveal qualifications not obvious from the main text.

If you are still developing an efficient reading routine, focus first on the high-information parts of the paper before trying to inspect every detail at once.

Separate statistical significance from practical importance

Large studies can detect very small differences. Whether those differences matter is a separate question.

Ask:

If this estimate is approximately correct, would it change a scientific explanation, clinical decision, policy, engineering choice, educational intervention, or other real decision?

The answer depends on context. There is no universal effect-size threshold that makes a result important.

Check whether the authors are making causal claims

Words matter.

Statements such as “X caused Y,” “X improved Y,” or “X reduced Y” imply something stronger than “X was associated with Y.”

When causal language appears, ask whether the study design and assumptions can support it.

Watch Out

Do not allow causal wording to outrun the design. An association can remain after statistical adjustment without proving that changing the exposure would change the outcome. Evaluate temporality, confounding, selection, measurement, and the design's causal assumptions before accepting causal interpretation.

Compare the conclusion with the results, not with the title

Titles and abstracts are highly compressed. Conclusions can also become more confident than the evidence warrants.

After reading the results, ask:

  • Does the conclusion answer the original research question?
  • Does it acknowledge uncertainty?
  • Does it distinguish association from causation?
  • Does it generalize beyond the studied population?
  • Does it emphasize a secondary or subgroup result more than the primary result?
  • Does it imply practical importance that was never demonstrated?
  • Does it acknowledge important contradictory findings?

A useful appraisal often consists of rewriting the conclusion in language proportional to the evidence.

Look for selective reporting

Researchers make many analytical decisions: which outcomes to emphasize, which time points to report, which covariates to include, which subgroups to show, and which statistical models to present.

If these choices are made after examining the results and only favorable analyses are highlighted, the published picture can be misleading.

When available, compare the paper with its preregistration, trial registration, protocol, or statistical analysis plan. Check whether prespecified primary outcomes and analyses match what appears in the publication.

Cochrane includes selection of the reported result as an explicit risk-of-bias domain for randomized trials.

Read the limitations, but conduct your own appraisal

Authors' limitation sections are useful because they show which problems the researchers themselves recognize.

But critical appraisal does not mean copying that paragraph into your notes.

Look for limitations the authors may not emphasize, and distinguish limitations that merely exist from limitations that materially weaken the conclusion.

A paper does not become invalid because it has limitations. Every study has constraints. The question is what those constraints mean for the claim you want to take from it.

Check conflicts of interest and funding

Read the funding statement and conflict-of-interest disclosures.

A declared financial relationship does not automatically invalidate a study, and absence of commercial funding does not guarantee unbiased research.

Instead, ask whether relationships could plausibly influence study design, comparison choice, data access, analysis, interpretation, reporting, or publication. Then consider the safeguards and transparency provided.

The International Committee of Medical Journal Editors emphasizes disclosure of relationships and activities that readers could perceive as influencing the work.

Check whether the reporting is complete enough to evaluate

Poor reporting and poor research are not identical, but incomplete reporting can make trustworthy appraisal impossible.

Reporting guidelines can help you determine what information should normally be available for particular study designs. The EQUATOR Network maintains a searchable library of health-research reporting guidelines, including CONSORT for randomized trials, STROBE for observational research, PRISMA for systematic reviews, and many design- or topic-specific extensions.

These guidelines primarily improve reporting rather than serving as simple quality scores. A paper can mention every checklist item and still contain methodological problems. Conversely, missing reporting may leave you uncertain about what was actually done.

Consider whether the findings apply beyond the study

Internal validity asks whether the study provides a credible answer within the conditions examined. External validity or applicability asks whether that answer is useful elsewhere.

Consider:

  • How similar is your target population to the study population?
  • Would the intervention, exposure, or setting operate similarly?
  • Are relevant demographic, institutional, cultural, geographic, or temporal differences important?
  • Were participants unusually selected?
  • Would implementation conditions differ?

A study can be internally strong but only narrowly applicable. That is not necessarily a flaw. It means you should state the boundaries of the evidence correctly.

Place the paper in the wider evidence base

One paper rarely settles an important scientific question.

Ask whether the finding is consistent with previous research, whether credible studies disagree, whether the paper replicates earlier evidence, and whether systematic reviews or other syntheses provide a broader picture.

Do not reject a surprising result simply because it disagrees with previous research. Equally, do not overturn an established evidence base because one new paper has a dramatic conclusion.

Use individual studies as pieces of evidence whose importance depends partly on their design, validity, precision, directness, and relationship to other evidence.

Do not confuse journal prestige with study validity

A journal's reputation, indexing status, Impact Factor, or quartile can tell you something about the publication venue. None substitutes for evaluating the study itself.

Peer review can improve manuscripts and filter some problems, but publication does not certify every analysis or conclusion as correct.

Read the paper, not merely the journal name.

04 · A Practical Example

From an impressive headline to a defensible interpretation

Hypothetical Example

A study reports that using an AI study tool improves university grades

Imagine a published observational study of 8,000 university students. Students who report using an AI study tool have higher average grades than students who report not using it. The paper concludes that AI-assisted studying improves academic performance.

1. Ask what the design can establish The study is observational. Students were not randomly assigned to use the tool, so the groups may differ in ways related to academic performance.
2. Examine selection and measurement Participation was voluntary, AI use was self-reported, and academic performance was measured using student-reported grade categories rather than verified institutional records.
3. Look for confounding The analysis adjusts for age and field of study but does not measure prior academic performance, study time, motivation, digital skills, or access to paid educational resources.
4. Examine the result The adjusted association remains statistically significant, but the estimated difference is small. The large sample makes even modest differences statistically detectable.
5. Compare the evidence with the conclusion The data support an association between reported AI-tool use and reported academic performance under the study's analysis. They do not, by themselves, demonstrate that giving students the tool would cause their grades to improve.
6. Write the defensible interpretation The study provides evidence of an association worth investigating further, but self-selection, measurement limitations, and residual confounding make a strong causal conclusion premature.

The paper is not therefore “useless” or necessarily “bad.” Critical appraisal changes the strength and wording of the conclusion you are justified in taking from it.

05 · What Researchers Often Get Wrong

Common mistakes in critical appraisal

Misconception

Critical appraisal means finding as many weaknesses as possible

A long list of minor limitations is less useful than identifying the few issues capable of materially changing the study's result or interpretation. Appraisal should assess consequences, not reward fault-finding.

Misconception

A large sample makes a study trustworthy

A large sample can improve precision, but it does not automatically correct confounding, selection bias, measurement error, inappropriate analysis, or other systematic problems. A precisely estimated biased result is still biased.

Misconception

A statistically significant result proves the hypothesis

Statistical significance does not establish that an effect is large, practically important, unbiased, causal, or true under every plausible assumption. Interpret the estimate, uncertainty, study design, and potential biases together.

Misconception

If the authors acknowledge the limitations, the problem is solved

Disclosure is valuable, but writing about a limitation does not remove its effect. You still need to determine whether the limitation materially weakens the conclusion.

Misconception

A checklist score tells me whether the paper is good

Design-specific appraisal tools can structure your thinking, but critical appraisal is not simply arithmetic. Different problems have different consequences, and some tools explicitly discourage reducing complex risk-of-bias judgments to an undifferentiated numerical quality score.

Misconception

A prestigious journal means the methods have already been verified for me

Peer review is not a substitute for your own evaluation. Published papers can contain methodological weaknesses, analytical errors, incomplete reporting, and conclusions open to legitimate disagreement.

06 · What This Means for You

Judge what the paper allows you to say

You rarely need to classify an entire paper as simply trustworthy or untrustworthy. A more useful outcome is to determine which claims you can reasonably take from it and with what level of confidence.

A simple appraisal framework

If the research question is unclear
Clarify what the study actually tested before interpreting its results.
If the design cannot support the strength of the authors' claim
Use more cautious language that reflects what the design can establish.
If an important source of bias could substantially change the result
Treat the finding with greater uncertainty rather than merely mentioning the bias in passing.
If the effect is statistically significant but very small
Ask whether the magnitude is scientifically, clinically, or practically meaningful.
If the study is methodologically strong but narrowly sampled
Separate confidence in the result for the studied population from confidence in applying it elsewhere.
If one study contradicts a much larger evidence base
Investigate why it differs and assess its methods before either dismissing it or treating it as decisive.

A useful final question is: If I cite this paper, what is the strongest sentence I can write that the study genuinely supports?

If your planned sentence is stronger than the design, measurements, results, and uncertainty permit, weaken the sentence rather than strengthening the evidence in your imagination.

07 · A Quick Checklist

Critical appraisal checklist for research papers

Before relying on a research paper, check:
Identify the exact research question, objectives, or hypotheses the study was designed to address.
Check whether the study design is appropriate for the question and the strength of inference being claimed.
Examine how participants, cases, observations, or data were selected and whether selection could bias the findings.
Check whether important exposures, interventions, outcomes, and other variables were defined and measured appropriately.
Identify plausible sources of bias, confounding, missing data, and measurement error and consider how they could affect the result.
Examine whether the statistical or qualitative analysis matches the design, data, and research question.
Read the actual effect estimates, uncertainty measures, tables, and figures rather than relying only on p-values or the abstract.
Compare the authors' conclusions with what the design and results can genuinely support, especially when causal language is used.
Review limitations, preregistration or protocols where relevant, funding, conflicts of interest, and signs of selective reporting.
Consider whether the findings apply to your population or context and how they fit with the wider body of evidence.
08 · Frequently Asked Questions

Frequently asked questions about critically evaluating research

What is critical appraisal in research?

Critical appraisal is the structured evaluation of research evidence to determine how trustworthy, meaningful, and applicable its findings are. It examines the relationship among the research question, design, conduct, analysis, results, and interpretation rather than merely looking for flaws.

What should I look at first when evaluating a research paper?

Start with the research question and study design. Until you understand what the researchers wanted to determine and how they attempted to determine it, you cannot judge whether the methods or conclusions are appropriate.

How can I tell whether a research paper is reliable?

There is no single reliability test. Examine the appropriateness of the design, participant or data selection, measurement quality, risk of bias, confounding, missing data, analysis, precision, transparency, and consistency between the evidence and conclusions. Use a design-specific appraisal tool when appropriate.

Does a peer-reviewed paper still need critical appraisal?

Yes. Peer review provides an important editorial evaluation but does not guarantee that a study is free from bias, analytical problems, incomplete reporting, or debatable interpretation. Readers remain responsible for evaluating evidence before relying on it.

Can I use a checklist to evaluate any research paper?

Use a tool appropriate to the study design. CASP and JBI, for example, provide different appraisal instruments for different forms of evidence because the threats to validity differ among randomized trials, observational studies, qualitative research, systematic reviews, and other designs.

What is the difference between reporting quality and research quality?

Reporting quality concerns whether the paper clearly describes what was done and found. Methodological quality or risk of bias concerns whether the study's design and conduct could produce a trustworthy result. Complete reporting helps you evaluate research, but a completely reported study can still have methodological weaknesses.

Is a statistically significant result more trustworthy?

Not necessarily. Statistical significance does not remove bias or confounding and does not tell you whether the effect is large or practically important. Evaluate the estimate, uncertainty, methodology, and potential sources of systematic error together.

Should I reject a paper if it has several limitations?

Not automatically. All studies have limitations. Ask whether the particular limitations materially threaten the claim you need from the study. A paper may provide useful evidence for a narrow conclusion even when it cannot support a broader one.

09 · The Bottom Line

Ask whether the evidence earns the conclusion

The Bottom Line

Critical evaluation is about tracing the logic of a study from its question through its design, measurements, analysis, and results to the conclusion.

Look for the sources of bias and uncertainty that could materially change what the findings mean, rather than simply collecting minor criticisms.

The most useful outcome is not a verdict that a paper is “good” or “bad.” It is a defensible understanding of what the study tells you, what it does not tell you, how confident you should be, and how strongly you can use it in your own research.

10 · Sources and Further Reading

Sources and further reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.