Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

When Should You Say the Evidence Does Not Generalize Rather Than Assuming It Does?

Evidence should not be assumed to generalize when differences between the available studies and the target population or setting could plausibly change the finding in an important way. The key is to identify what differs and why that difference matters.

574
When Evidence Does Not Generalize Guide 574 of 899
01 · The Question

When does extrapolation become too difficult to defend?

Almost every use of research involves some generalization. The participants in a study are not identical to every person who may later receive the intervention, encounter the exposure, or make a decision based on the findings.

Usually, that is not a fatal problem. Many population differences do not substantially change relative effects. But sometimes the available evidence comes from people, interventions, comparators, outcomes, or settings that differ from your target question in ways that could matter.

At what point should you stop saying, “The evidence probably applies,” and instead acknowledge that generalization is uncertain or unsupported?

02 · The Short Answer

Do not assume generalizability when effect-relevant differences are substantial

In Brief

You should not assume that evidence generalizes when the studied population or context differs from your target in ways that create a credible possibility of a substantially different effect, outcome, or interpretation.

The existence of any difference is not enough. Identify the mismatch, explain why it could matter, examine any evidence about effect modification, and calibrate your conclusion to the uncertainty. Sometimes the appropriate conclusion is not “it does not generalize,” but the more precise “we do not know whether it generalizes.”

03 · What You Need to Know

Generalizability is not an all-or-nothing property of a study

A study does not carry a permanent label saying “generalizable” or “not generalizable.” Its applicability depends on the target question.

Evidence from one population may be highly relevant to another population for one outcome but less relevant for another. A treatment's adverse effects might transfer reasonably well across related conditions even when its benefits do not. Likewise, an intervention may retain its biological mechanism across settings while its real-world effectiveness changes because implementation differs.

Core GRADE frames this issue as indirectness. Researchers specify the target population, intervention, comparator, and outcome, then examine how closely the available evidence corresponds to that target. When a substantial mismatch creates a credible likelihood that the magnitude of effect will differ substantially, concern about indirectness increases.

Start by defining what you want to generalize to

“Does this study generalize?” is incomplete. Generalize to whom, where, for which intervention, compared with what, and for which outcome?

Suppose a trial enrolled adults aged 40 to 60. Whether its findings generalize depends on your target. Adults aged 35 to 65 might present relatively little concern for some questions. Infants, very old adults, or people with a condition fundamentally different from those studied might present much more.

This is why you should first define the target population and question. Only then can you identify meaningful mismatches.

A difference matters when it could change the effect

Demographic mismatch alone should not automatically trigger a claim of non-generalizability. Core GRADE notes that population differences encountered within otherwise direct evidence often do not justify downgrading because true differences in relative effects across many patient characteristics are uncommon. The concern becomes stronger when there is a credible reason to expect the mismatch to alter the effect.

Age, for example, may matter profoundly for an intervention dependent on developmental physiology but much less for another intervention within a restricted adult age range. Geography may matter when it represents differences in infrastructure, epidemiology, institutions, or implementation, but a national border by itself is not a mechanism.

The useful question is therefore not “Are these populations different?” It is “Are they different on something capable of changing the finding?”

Visible difference The study and target populations differ on a characteristic such as age, location, diagnosis, or socioeconomic composition.
Effect-relevant difference There is a credible reason that the characteristic could alter the magnitude, direction, mechanism, or practical interpretation of the finding.

Look beyond the population

Generalizability problems are often described as population problems, but the mismatch may lie elsewhere.

The target intervention may differ in dose, intensity, duration, personnel, or implementation. The comparator may represent a different standard of care. The study may measure a surrogate when the target question concerns a patient-important outcome. Follow-up may be too short for the decision being made.

GRADE treats these population, intervention, comparator, and outcome mismatches as possible sources of indirectness rather than focusing exclusively on participant demographics.

A population can therefore look remarkably similar to yours while the evidence remains indirect for the actual decision.

Setting can make an otherwise transferable intervention behave differently

Some interventions depend strongly on context. Staffing, infrastructure, institutional rules, cultural practices, technology, service availability, implementation capacity, or baseline alternatives may influence what participants actually receive.

An intervention demonstrated in specialist centers may not produce the same result when implemented without specialist personnel. A digital program requiring reliable connectivity may operate differently where access is intermittent. An educational intervention built around small classes may change when transferred to classrooms containing substantially more students.

When these conditions are central to how an intervention produces its effect, they should not be treated as incidental background characteristics.

The question then becomes whether evidence from the other population is sufficiently relevant to yours, rather than whether the populations merely resemble one another.

Do not mistake absence of subgroup evidence for proof of generalizability

Suppose a meta-analysis finds no statistically significant difference between younger and older participants. That may support transferability, but it does not automatically prove equal effects.

Subgroup analyses often have limited power. Published reports may provide inadequate subgroup data. Study-level comparisons may also be confounded or affected by ecological bias.

Cochrane recommends cautious interpretation of subgroup analyses and emphasizes that credible effect modification should preferably be based on prespecified hypotheses, direct comparisons between subgroup effects, clinical plausibility, and, where possible, replicated within-study relationships.

“No demonstrated difference” and “demonstrated similarity” are not the same claim.

Do not mistake uncertainty for evidence of failure to generalize

The opposite error is equally important. If the target population was not adequately studied, you may not know whether the effect differs.

That does not necessarily justify saying, “The intervention does not work in this population.” The evidence may simply be insufficient to establish whether it does.

Evidence suggests the finding differs Empirical or strongly supported mechanistic evidence indicates that the effect or finding changes in the target population.
Generalizability is uncertain The target population is insufficiently represented or important differences exist, but the available evidence cannot establish whether the finding changes.

That distinction is not semantic housekeeping. It prevents absence of evidence from quietly becoming evidence of absence, an old statistical trick that refuses to retire.

Consider absolute as well as relative effects

Evidence can generalize differently depending on which effect you mean.

A relative treatment effect might remain approximately stable across populations while absolute benefits differ because baseline risk changes. For example, the same relative risk reduction can produce a much larger absolute benefit in a high-risk population than in a low-risk population.

It would therefore be misleading to say simply that the “effect generalizes” without specifying whether you mean the relative effect, absolute effect, implementation outcome, or another quantity.

When mechanisms change, concern becomes stronger

A particularly important warning sign appears when the target population lacks a condition required for the proposed mechanism to operate, or possesses a characteristic expected to alter it substantially.

Core GRADE provides instructive examples in which changes to the target population can make earlier evidence highly indirect. When important biological characteristics of the target change, such as viral variants that undermine the mechanism of an antibody treatment, earlier trials may no longer provide direct evidence for the new population.

The same reasoning applies beyond clinical research. If a program requires institutional autonomy, literacy, technology, or another prerequisite that the target population or setting lacks, the transferability question becomes more serious.

Sometimes the evidence should be narrowed rather than discarded

A failure to generalize to one population does not invalidate the original finding. It limits its domain.

If an intervention has strong evidence among adults but uncertain applicability to children, the appropriate conclusion may remain strong for adults while explicitly limited for children. If an educational program has been demonstrated only in highly resourced schools, its evidence can remain persuasive for comparable schools while being uncertain elsewhere.

This preserves what the evidence genuinely establishes without extending it beyond its evidential reach.

04 · A Practical Example

When a successful intervention should not simply be assumed to transfer

Hypothetical Example

A remote mental-health program moved into a different service environment

Imagine a hypothetical evidence base showing that a remote psychological intervention improves outcomes among adults treated through specialist mental-health services.

The original evidence Participants receive the digital program alongside professional assessment, scheduled clinician monitoring, reliable internet access, and rapid referral when symptoms worsen.
The target population A community program proposes using the same digital materials among people with similar symptoms but limited access to clinicians and inconsistent connectivity.
The superficial argument The participants have similar symptom profiles, so the evidence should generalize.
The important mismatch Professional monitoring and referral may be part of how the intervention operates safely and effectively, while connectivity affects actual exposure to the program.
The defensible conclusion The original evidence supports the intervention under the studied service conditions. Whether the same effect occurs in the target setting is uncertain because effect-relevant implementation conditions differ substantially.

Notice what the conclusion does not say. It does not claim that the intervention will fail. It says the available evidence does not justify assuming that the original effect transfers unchanged.

05 · What Researchers Often Get Wrong

Common mistakes when judging generalizability

Misconception

Any population difference means the evidence does not generalize

Most study and target populations differ somehow. The relevant issue is whether those differences plausibly change the finding in an important way. Core GRADE explicitly bases indirectness judgments on the likelihood of substantial differences in effects.

Misconception

If no subgroup difference is statistically significant, generalizability is established

Subgroup analyses can be underpowered, poorly reported, or confounded. Failure to detect effect modification does not automatically establish equivalent effects across populations.

Misconception

If my population was not studied, the intervention does not work for them

Lack of direct evidence creates uncertainty; it does not necessarily demonstrate absence of effect. Distinguish evidence of a meaningful difference from insufficient evidence about whether a difference exists.

Misconception

Demographic similarity guarantees applicability

Similar participants can receive different interventions, comparators, implementation support, or outcome assessments. Applicability depends on the whole research question, not demographic resemblance alone.

Misconception

A lack of generalizability makes the original study weak

A rigorous study can answer its own question convincingly while providing indirect evidence for a different question. Internal credibility and applicability to another target are distinct issues.

06 · What This Means for You

Use the strongest conclusion the evidence can support, but no stronger

When assessing generalizability, specify the target first and then compare it with the evidence. Focus on characteristics that could plausibly modify the finding rather than compiling every observable difference.

A simple decision framework

If differences between the study and target populations are unlikely to alter the effect
Generalization may be reasonable despite demographic or geographical differences.
If an important effect modifier differs substantially between the evidence and target population
Do not assume the original effect transfers unchanged; explicitly consider indirectness.
If empirical evidence shows a credible effect difference across the relevant populations
Report the population-specific difference rather than presenting the overall effect as universally applicable.
If the target population is poorly represented and evidence about effect modification is inadequate
Say that generalizability is uncertain rather than claiming either equivalence or failure.
If the evidence applies well to one population but poorly to another
Narrow the conclusion to the population actually supported instead of discarding the entire evidence base.

This distinction becomes particularly important when evidence comes from settings with substantially different resources. What looks like a population mismatch may actually be an implementation or contextual mismatch.

Where evidence consistently survives such contextual differences, the opposite inference may become relevant: consistency across substantially different settings can strengthen confidence that the finding is not restricted to one narrow environment.

07 · A Quick Checklist

Before saying that evidence generalizes, check:

Before extrapolating to your target population, check:
Have you explicitly defined the population and setting to which you want to generalize?
Which differences between the study and target populations could plausibly modify the finding?
Is there empirical or mechanistic evidence that those characteristics matter?
Are the intervention and comparator sufficiently similar to those relevant to your target?
Are the outcomes and follow-up periods appropriate for the target question?
Could implementation or contextual differences change how the intervention operates?
Are you distinguishing absence of evidence for a difference from evidence that no difference exists?
Have you distinguished relative effects from potentially different absolute effects?
Is your conclusion appropriately limited when important uncertainty remains?
08 · Frequently Asked Questions

Questions about limits to generalizability

How different must a population be before evidence no longer generalizes?

There is no universal numerical threshold. The important question is whether the mismatch creates a credible likelihood that the magnitude or interpretation of the effect will differ substantially.

Does evidence from another country automatically have limited generalizability?

No. Country matters when it represents contextual characteristics capable of altering the finding. Evidence from another country may be highly relevant when the effect-relevant conditions resemble those in your target setting.

What should I say when the target population was not adequately studied?

If there is insufficient evidence to determine whether the effect differs, state that applicability or generalizability is uncertain. Do not convert inadequate evidence into a claim that the intervention either definitely works or definitely fails in that population.

Can strong evidence still be indirect?

Yes. A study may provide highly credible evidence for the population and conditions actually studied while offering less direct evidence for a different target question. GRADE treats this mismatch as indirectness rather than as a defect in the original study.

Can relative effects generalize while absolute effects do not?

Yes. Different baseline risks can produce different absolute benefits or harms even when relative effects are similar. Specify which effect measure is relevant to the decision.

Should I exclude indirect evidence entirely?

Not necessarily. Core GRADE explicitly recognizes that indirect evidence can be valuable, particularly when direct evidence is unavailable or uncertain. Its indirectness should be evaluated rather than the evidence being automatically discarded.

09 · The Bottom Line

Do not generalize farther than the evidence can reasonably travel

The Bottom Line

Do not assume that evidence generalizes when important differences between the studied and target populations or settings create a credible possibility that the effect or its interpretation will change substantially.

Focus on effect-relevant differences rather than superficial mismatch, and distinguish demonstrated non-generalizability from uncertainty about generalizability. Often the scientifically defensible response is not to reject the evidence, but to narrow the claim to the populations and conditions it can actually support.

10 · Sources and Further Reading

Sources and further reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes