03 · What You Need to Know
Judge Context by Its Relationship to the Claim You Are Testing
There is no percentage of difference that triggers a new study
Researchers sometimes look for an implicit threshold: must the new population be 20% different, come from another country, use another language, or operate under a different institutional system? Research methodology offers no general cutoff of this kind.
External validity concerns whether inferences supported in one study can reasonably extend to other populations or circumstances. That question depends on what is being generalized and what characteristics might influence it. Consequently, contextual similarity and difference need to be evaluated relative to the particular claim rather than against a universal standard.
A useful starting point is to specify the target of generalization. Are you asking whether an association exists, whether an intervention causes an outcome, whether the magnitude of that effect remains similar, whether a measurement instrument functions equivalently, or whether an intervention can be implemented successfully? Different claims make different contextual characteristics relevant.
Not every difference is an effect modifier
Suppose two populations differ in average age, income, language, class size, technology access, institutional policy, and cultural background. Listing these differences establishes that the populations are not identical. It does not establish that the previous finding should change.
The stronger argument identifies a characteristic that could function as an effect modifier, boundary condition, implementation constraint, or source of measurement non-equivalence. In causal research, effect modification means that the effect of an exposure or intervention varies according to another characteristic. More broadly, a contextual factor may matter because it changes how a mechanism operates.
This distinction prevents researchers from treating context as a decorative paragraph in the introduction. If you argue that a particular difference justifies the study, your research design should ideally measure that difference and allow you to investigate its relevance.
Descriptive difference
The new setting differs from previous settings, but there is no developed reason to expect the difference to affect the finding.
Scientifically meaningful difference
The difference is plausibly connected to the mechanism, measurement, outcome, implementation, or inference being investigated.
The distinction is central to deciding whether a contextual difference is scientifically meaningful.
Ask whether the relevant mechanism could operate differently
A mechanism explains how or why an exposure, intervention, or condition produces an outcome. Context becomes particularly important when it can change that mechanism.
Consider a digital learning intervention that improves achievement partly because students receive immediate individualized feedback and teachers use the resulting data to adjust instruction. If the proposed local setting has intermittent internet access, shared devices, and little teacher preparation time, these conditions are directly connected to how the intervention is supposed to produce its effect.
Contrast that with a difference that is easy to observe but unrelated to the proposed mechanism. The fact that the new schools are located in another province, for example, adds little by itself. Geography becomes scientifically relevant only when it represents characteristics that matter to the process under investigation.
Population composition can matter without making every demographic difference important
Age, socioeconomic status, prior knowledge, baseline risk, language, disease burden, educational attainment, and other population characteristics may affect outcomes or intervention responses. Whether they justify another study depends on their relevance to the particular phenomenon.
A population difference becomes more persuasive when previous evidence suggests heterogeneous effects, theory predicts different responses, or the characteristic is closely related to the causal process. If previous studies already include substantial variation in that characteristic and findings remain stable, the argument for another study based solely on that difference may be weaker.
Cultural differences require a mechanism, not an adjective
Culture can influence behavior, social expectations, response styles, interpretations of constructs, intervention acceptability, and interactions with institutions. These possibilities can make cultural differences relevant to a new study.
But “the culture is different” remains too broad. Which cultural characteristic differs? What part of the research question could it influence? Is there prior evidence or theory supporting that expectation? Would the measurement instrument retain the same meaning? Could the intervention encounter different norms or incentives?
The more precisely those questions can be answered, the stronger the contextual rationale becomes.
Institutional differences may change the environment in which a finding occurs
Healthcare systems, educational structures, regulatory environments, organizational procedures, labor arrangements, and public policies can change how interventions or exposures operate. An educational program dependent on teacher autonomy may function differently where curricula and instructional schedules are centrally prescribed. A healthcare intervention requiring rapid specialist referral may encounter different constraints in a system where referral pathways are limited.
Such differences can justify additional research when the institutional feature is integral to the mechanism or implementation. The fact that institutions have different names or administrative structures is less important than whether those structures alter what participants can actually do.
For this reason, the case for research based on differences in healthcare, education, policy, or institutions should specify the operational consequence of the difference.
Resource differences matter when the phenomenon depends on those resources
Funding, staffing, infrastructure, equipment, time, expertise, internet access, transportation, and other resources can affect feasibility, fidelity, exposure, participation, and outcomes. An intervention evaluated with intensive professional support may not behave identically when implemented through ordinary services with fewer personnel.
Resource differences are particularly persuasive when they affect an ingredient that existing evidence suggests is necessary for the intervention or process to work. Researchers considering whether resource differences justify a new study should therefore connect resources to expected mechanisms rather than simply describing one setting as “resource constrained.”
Look at how broad the existing evidence already is
Contextual difference should not be evaluated by comparing the proposed setting with a single convenient study when a much larger evidence base exists.
Suppose a relationship has been observed across dozens of studies involving different age groups, socioeconomic conditions, countries, institutions, and implementation arrangements. That heterogeneity may already provide evidence that the finding survives substantial contextual variation. Another setting that falls within the range already represented may add relatively little.
The situation is different when nearly all existing evidence comes from a narrow population or highly standardized environment. A new setting that differs along a theoretically relevant dimension can then provide an informative test of generalizability.
This is one reason replication can be scientifically productive. Nosek and Errington argue that a replication is informative when possible outcomes provide diagnostic evidence about an existing claim. A replication that succeeds under meaningfully different conditions can extend confidence in generalizability, while a failure can reveal previously unrecognized constraints on the claim.
Ask whether both possible outcomes would teach you something
Before declaring the context different enough, imagine that your proposed study produces the same result as previous research. What would you conclude? Then imagine that it produces a different result. What would that tell you?
If a consistent result would demonstrate that the finding survives an important contextual change, and an inconsistent result would provide credible evidence of a boundary condition, the study has a strong inferential purpose.
If a consistent result would merely produce another local estimate and an inconsistent result would be difficult to interpret because no contextual mechanism was specified, the rationale is weaker.
Watch Out
Do not decide that a context is scientifically different simply because you can write a long list of differences between it and previous settings. Almost any two populations can be made to look dramatically different when enough characteristics are listed. Scientific relevance depends on which differences bear on the inference.
The necessary study may not be a full replication
Even when a meaningful contextual difference exists, repeating the entire previous study is not automatically the best design. If uncertainty concerns measurement, a validation or measurement-invariance study may be more appropriate. If the issue is implementation, implementation research may answer the question more directly. If local baseline risk or prevalence is missing, descriptive or surveillance data may be sufficient.
The broader principle is that local research becomes necessary when existing evidence leaves an important local inference unresolved. Your design should target that unresolved inference rather than treating local data collection as an end in itself.