01 · The Question
Who Do the Authors Say Their Findings Apply To?
A paper reports findings from the people, organizations, schools, communities, records, or other units that actually entered the study. Yet the conclusion may speak about a much larger population. That move from the observed sample to a broader population deserves scrutiny.
Suppose researchers survey students from two universities but conclude that their findings describe university students nationally. Or a clinical study recruits participants from specialist hospitals but discusses implications for all patients with the condition. The issue is not simply whether the sample is large. The more fundamental question is whether the evidence collected from that sample can reasonably support the population-level claim being made.
Evaluating this requires you to identify both sides of the inference: who was actually studied and who the authors want you to believe the findings apply to.
03 · What You Need to Know
Compare the Sample With the Population Behind the Claim
Start by identifying the target population
You cannot judge whether a sample is appropriate without first asking: appropriate for whom? Generalizability is meaningful only relative to a specified target population. A sample might support an inference about one population while providing much weaker evidence about another.
Look at the research question, eligibility criteria, methods, abstract, discussion, and conclusion. Authors do not always state the target population explicitly, so pay attention to the nouns they use when interpreting their results. Do they write about “participants in this study,” “students at the participating universities,” “Filipino university students,” or simply “university students”? Those phrases imply progressively broader populations.
Study sample
The units that actually contributed data to the analysis.
Target population
The population to which the authors intend the relevant findings or estimates to apply.
This distinction is easy to overlook because papers sometimes move quietly from describing their participants to making broader statements. Your task as a critical reader is to make that inferential jump visible.
Trace how people moved from the target population into the sample
Next, reconstruct the selection process. Who could theoretically belong to the target population? Who was accessible to the researchers? Who was eligible? Who was invited or approached? Who agreed to participate? Who ultimately contributed usable data?
Each stage can change the composition of the sample. Geography, institutional access, recruitment channels, language requirements, internet access, clinical eligibility criteria, willingness to participate, and many other mechanisms may make some people more likely to enter a study than others.
This is why statements such as “we recruited 2,000 participants” tell you surprisingly little by themselves. You also need to know where those 2,000 came from. A very large sample drawn through a restrictive recruitment process may still differ systematically from the population about which the authors make claims. The distinction between sample size and sample selection becomes especially important when comparing a large convenience sample with a smaller, deliberately selected sample.
Do not reduce appropriateness to demographic resemblance
A common appraisal strategy is to compare age, sex, socioeconomic status, ethnicity, educational level, disease severity, or other observable characteristics between the sample and target population. Such comparisons can be informative, but visual resemblance on a demographic table is not a complete test of external validity.
The crucial question is whether characteristics related to selection also matter for the outcome, association, or effect being studied. For causal effects in particular, differences in effect modifiers between the study sample and target population can cause an effect estimated in the sample to differ from the effect that would be expected in the target population. Consequently, some sample-population differences may matter considerably while others may have little relevance to the particular inference.
This is also why asking whether a sample is simply “representative” can be too vague. Representativeness should be tied to a specified population, a particular result, and assumptions about how that result extends beyond the observed sample. Whether a representative sample is necessary for a particular research question therefore needs to be considered separately rather than treated as a universal requirement.
Descriptive estimates often place heavier demands on sampling
Consider a study estimating the prevalence of academic burnout among all students at a university. If students experiencing severe burnout are much less likely to participate, the observed prevalence may not accurately describe the university population. Here, the composition of the sample is directly connected to the quantity researchers are trying to estimate.
Questions about prevalence, proportions, averages, voting intentions, population attitudes, and other population characteristics often make the sample-to-population relationship particularly important. A percentage calculated from a sample is mathematically precise for the observed data, but that precision does not automatically make it a valid population estimate.
Analytical and causal questions require a somewhat different judgment
Not every study is trying to estimate a population prevalence or mean. Some investigate associations, mechanisms, or causal effects. For these questions, lack of proportional demographic representation does not automatically invalidate the study.
For example, a randomized trial may provide a credible treatment-effect estimate for its enrolled participants even though those participants do not mirror every characteristic of the wider patient population. The separate question is whether that effect can then be generalized to the population of interest. Internal validity and external validity are therefore related but distinct concerns.
This distinction prevents an overly mechanical appraisal. A nonrepresentative sample is not automatically useless, just as a superficially representative sample is not automatically adequate. Even a convenience sample can provide useful evidence for some research questions, provided the inference does not outrun what the design can support.
Look for consequential undercoverage rather than demanding perfect similarity
Ask whether segments of the target population had little or no realistic opportunity to enter the study. A web-based survey may underreach people with limited internet access. Recruitment through a single professional association may miss professionals outside that network. A study conducted during working hours may disproportionately miss people whose employment prevents participation.
The seriousness of such undercoverage depends on what is being estimated. If the excluded or underrepresented groups would plausibly have different values of the outcome, different exposure patterns, or different responses to an intervention, extending the sample result to them requires stronger justification.
Sometimes the problem is more explicit: entire subpopulations were excluded by design. In that situation, examine whether important groups were excluded from the study and whether the authors nevertheless include those groups within their conclusions.
Eligibility criteria define the population the study can directly speak about
Eligibility criteria are not inherently a weakness. Researchers often need them for safety, measurement quality, feasibility, or to answer a deliberately focused question. The problem arises when a narrowly eligible sample is paired with conclusions phrased as though a substantially broader population had been studied.
For example, evidence from adults aged 18–35 without major comorbidities may be highly informative about that defined population. It does not automatically establish the same effect among older adults or people with substantial comorbidity. When reviewing such a paper, ask whether the eligibility criteria make the sample narrower than the conclusion.
Participation after recruitment matters too
An apparently strong sampling frame can still produce a selectively participating sample. Suppose researchers randomly invite people from a well-defined population but only a small fraction participate. Random selection at the invitation stage does not guarantee that respondents remain comparable with the intended population if participation itself is selective.
A low response or participation rate is not, by itself, a numerical threshold proving bias. What matters is who did and did not participate and whether participation is associated with variables relevant to the result. The implications of low participation for generalizability therefore depend on the mechanism of nonparticipation, not merely the percentage reported.
Ask whether statistical adjustment changes the inference
Researchers sometimes use weighting, standardization, outcome modeling, or related methods to generalize or transport estimates from a study sample to a defined target population. Such methods can be valuable when appropriate target-population information is available.
They are not magic repairs for an arbitrary sample. Their validity depends on assumptions, including adequate measurement of variables needed to account for relevant differences between the sample and target population. When authors use these methods, inspect what characteristics were measured, how the target-population data were obtained, what model was used, and which assumptions the analysis requires.
Watch Out
Do not judge sample appropriateness from sample size alone. Precision addresses sampling variability around an estimate; it does not by itself correct systematic differences between the people studied and the population to which the estimate is being extended.
06 · What This Means for You
Judge the Claim and the Sample Together
When critically reading a paper, resist the temptation to classify the sample as simply “good” or “bad.” Instead, evaluate the compatibility between the sample and the claim. The same sample can be entirely appropriate for one inference and inadequate for another.
A simple decision framework
If the conclusion describes only the population actually represented by the study design
Check whether recruitment and participation still introduce meaningful selection problems, but do not demand broader representativeness than the question requires.
If the authors estimate a population prevalence, proportion, or mean
Examine the sampling frame, selection probabilities, nonresponse, coverage, weighting, and relevant differences between the achieved sample and target population especially carefully.
If the study estimates an association or causal effect
Separate internal validity from external validity, then ask whether factors affecting selection could also alter the association or effect in the target population.
If important groups could not enter the sample
Check whether the conclusion nevertheless includes those groups and what evidence, if any, supports extending the findings to them.
If the sample and target population differ substantially
Look for substantive justification or appropriate generalization methods rather than assuming that statistical significance or a large sample solves the problem.
Pay particular attention to the wording of the conclusion. “Among participants in this study,” “among students at the participating institutions,” and “among university students nationally” are not interchangeable claims. Sometimes the most defensible correction to a study is not a different analysis but a narrower sentence.
When the evidence cannot support the broader population, the scientifically appropriate response may be to limit the inference to the population actually supported by the study. That is not a failure of the research. It is accurate calibration of what the evidence can tell us.
07 · A Quick Checklist
How to Check Whether the Sample Fits the Population Claim
When evaluating a study sample, check:
Identify exactly who contributed data to the final analysis.
Write down the population to which the authors explicitly or implicitly apply their findings.
Trace the selection process from the target population through eligibility, recruitment, participation, and final analysis.
Identify groups that were excluded, undercovered, unlikely to participate, or lost before analysis.
Ask whether sample-population differences concern variables relevant to the outcome, association, or treatment effect.
For population estimates, inspect the sampling design, response process, and any weighting or adjustment used.
Do not use sample size or statistical significance as substitutes for evaluating sample selection.
Compare the scope of the abstract and conclusion with the population the study design can reasonably support.