03 · What You Need to Know
Judge the gap against the inference you are trying to make
First define the target population
You cannot determine whether a population's absence threatens generalizability until you specify the population to which you want the findings to apply.
This is a basic principle in methodological work on external validity. Stuart and colleagues argue that discussion of generalizability or transportability is not meaningful without identifying the target population. The same study may support inference to one population more readily than another.
Suppose a study estimates the effect of an educational intervention among secondary-school students in a particular school system. There are several possible claims: the intervention worked for the students who participated; it works for students throughout that school system; it works for secondary-school students nationally; or it works for secondary-school students generally.
Those claims extend progressively beyond the evidence directly observed. A population missing from the sample may be irrelevant to the first claim and highly relevant to the fourth.
Evidence for the study sample
What the design and data support about the participants actually studied.
Evidence for the target population
What researchers want to infer about a broader or different population beyond the observed sample.
Strong internal validity does not automatically establish external validity
A well-designed randomized study can estimate an intervention effect credibly for its study sample while still leaving uncertainty about the corresponding effect in a broader target population. Generalizability research distinguishes the study sample average treatment effect from the target population average treatment effect because the two need not be equal when study participants do not adequately represent relevant features of the target population.
This is not a criticism of randomization. Randomization addresses treatment assignment within the study; it does not randomly assign people into the study from every population to which someone may later wish to apply the result.
A study can therefore be methodologically rigorous and still have a limited evidential scope.
The missing population matters most when relevant effect modifiers differ
For causal effects, one particularly important issue is effect modification. An effect modifier is a characteristic across which the effect of an intervention varies. If such characteristics have different distributions in the study sample and target population, the average effect estimated in the study may not represent the average effect in the target population.
Imagine that an intervention works especially well for participants with high prior knowledge. If the original studies disproportionately include such participants while the target population contains many people with low prior knowledge, applying the study's average effect directly to the target population may be misleading.
The same reasoning applies beyond a single demographic characteristic. Relevant differences might involve baseline risk, institutional resources, language, prior experience, environmental exposure, implementation support, socioeconomic conditions, or other features tied to the mechanism producing the outcome.
This is why demographic absence should not be treated mechanically. A missing characteristic matters because of what it may imply for the phenomenon, not merely because a column in the demographic table is empty.
You do not need proof of a different effect before confidence can be limited
There is an epistemic problem here. If researchers required prior proof that the effect differs before questioning generalizability, they would often need the very population-specific evidence whose absence created the uncertainty.
A more reasonable standard is plausibility supported by theory, previous evidence, mechanism, or relevant contextual differences. You can have less confidence that a finding applies without asserting that it definitely will not apply.
That distinction is central when considering whether you must expect a different finding before studying another population. Uncertainty and predicted difference are not equivalent.
Some findings are more portable than others
Not every research finding is equally sensitive to population composition.
A result tied closely to a particular institutional system, language, technology infrastructure, baseline risk profile, social environment, or implementation process may have more limited portability than a finding whose mechanism is well understood and supported across heterogeneous settings.
The existing evidence base also matters. One narrow study and twenty studies spanning diverse populations create different levels of confidence. If a finding has already replicated across settings that vary on the factors most plausibly relevant to the outcome, the absence of one additional population may create less uncertainty.
Conversely, repeated studies conducted in very similar samples do not necessarily provide the same external-validity evidence as studies spanning substantively different conditions. Twenty convenience samples can sometimes be twenty versions of the same inferential limitation.
Ask whether the population is outside the support of the existing evidence
Generalizability and transportability methods can sometimes use observed information about a study sample and target population to estimate effects beyond the original sample. These approaches rely on assumptions, including adequate overlap in relevant covariates between study and target populations.
If members of the target population have combinations of relevant characteristics that are essentially absent from the study sample, statistical adjustment cannot conjure empirical information where there is no meaningful overlap. Methodological discussions describe this through positivity or overlap conditions required for generalizing or transporting causal effects.
This provides a more precise way to think about some missing populations. The concern is not simply that a demographic label is absent. It is that the existing data may contain little or no empirical information about people with the characteristics needed to support the target inference.
Implementation can limit applicability even if the underlying effect is stable
Existing evidence may also lose relevance because the conditions required to produce the observed effect are different.
Suppose an intervention is effective when delivered with intensive staff training, reliable technology, small class sizes, or substantial institutional support. A target population operates in settings where those conditions are uncommon. Even if the intervention's underlying mechanism does not vary by population identity, the observed study effect may not describe what happens when the intervention is implemented under different conditions.
Research on external validity explicitly recognizes implementation, measurement, and setting differences as possible reasons effects observed in trials may not carry over straightforwardly to target populations.
This is one reason inclusion may matter without an expected different effect. Direct evidence can tell you whether the conditions necessary for the finding actually hold.
Measurement can create another boundary on confidence
Sometimes the concern is not whether the underlying phenomenon differs but whether it has been measured comparably.
If an instrument was developed and validated almost entirely within one population, researchers should consider whether its language, constructs, response options, administration procedures, or assumptions remain appropriate elsewhere. A score that looks comparable numerically may not necessarily represent the same construct in precisely the same way.
This does not mean every instrument requires complete redevelopment for every population. It means that confidence in substantive conclusions depends partly on confidence that the measurement supports the comparison or inference being made.
The consequences of extrapolation should affect how much evidence you demand
Suppose two research claims are equally uncertain. One informs a minor interface preference. The other informs access to an educational program, clinical treatment, or public service.
It would be strange to demand exactly the same evidential threshold for both.
The cost of being wrong should influence how cautiously evidence is extrapolated. When decisions are consequential, irreversible, difficult to monitor, or affect populations already poorly represented in the evidence base, direct evidence may become more important.
This does not transform methodological uncertainty into an ethical veto. It means that confidence is partly decision-relative: the amount of uncertainty tolerable for exploratory discussion may be inappropriate for a high-stakes policy claim.
Historical exclusion can make the evidence gap persistent rather than accidental
A missing population may be absent because of eligibility criteria, inaccessible recruitment procedures, institutional mistrust, language restrictions, geographic concentration of research sites, or other recurring features of research practice.
When exclusion is systematic, waiting for the evidence base to become representative on its own may simply reproduce the same gap. The rationale for research involving a historically excluded population can therefore include both inferential uncertainty and the processes that created it.
Still, history should not substitute for methodological reasoning. Researchers should specify which claims are uncertain and what direct evidence would improve confidence.
Watch Out
Do not turn uncertainty about generalizability into a claim that existing evidence is "invalid." Evidence can be internally credible for the population and conditions studied while remaining uncertain for another target population. State the boundary of the evidence rather than dismissing the evidence altogether.
04 · A Practical Example
When a missing population changes how confidently you can apply a finding
Hypothetical Example
An online intervention studied almost entirely in well-connected universities
Suppose several rigorous studies find that a synchronous online tutoring program improves student performance. Nearly all participating universities have reliable campus connectivity, extensive technical support, and students with routine access to suitable devices.
A university system wants to implement the program across remote campuses where connectivity is intermittent and device access is less consistent.
What the evidence establishes The tutoring program has produced beneficial outcomes under the conditions represented in the existing studies.
What is missing Students and institutions operating under substantially different access and infrastructure conditions are poorly represented.
Why the absence matters Reliable participation in synchronous tutoring is part of the pathway through which the intervention can produce its effect.
Appropriate conclusion Existing evidence supports the intervention under the studied conditions, but confidence about its effectiveness under substantially different access conditions is lower.
Next research question Can the intervention be implemented with sufficient participation and benefit in the target population, and which access conditions influence its effectiveness?
Notice what the conclusion does not say. It does not declare the previous studies wrong. It does not assume remote-campus students are inherently different. It identifies a condition relevant to the intervention and explains why the existing evidence provides limited information about that condition in the target population.