Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

When Should Research Move From “Does It Work?” to “For Whom, When, and Why?”

Once credible evidence supports an average effect, repeatedly asking only “Does it work?” may become less informative. Research can then examine who benefits, under what conditions effects vary, and what mechanisms might explain those differences.

768
When Research Should Move Beyond “Does It Work?” Guide 768 of 899
01 · The Question

When is the average effect no longer the most important question?

Early in the study of an intervention, program, treatment, or practice, asking whether it produces an effect is entirely reasonable. Researchers first need credible evidence that something happens at all.

As evidence accumulates, however, the average answer can become less informative. An intervention that produces a modest benefit on average might produce a substantial benefit for some people, little change for others, or different effects under different conditions.

At that point, continuing to ask only “Does it work?” can conceal the next scientifically important questions: for whom does it work, when does it work, under what conditions does it work, and why?

02 · The Short Answer

Move beyond the average effect when variation becomes the consequential uncertainty

In Brief

Research should increasingly move from “Does it work?” toward “For whom, when, and why?” when credible accumulated evidence has reasonably established the average effect and the more consequential uncertainty concerns variation in effects, boundary conditions, or the processes that produce them.

The transition is not a rigid stage that occurs after a fixed number of studies. Average effects, heterogeneity, moderators, and mechanisms can often be investigated together, and evidence about the overall effect should continue to be updated when relevant new information appears.

03 · What You Need to Know

An average effect can answer one question while hiding another

“Does it work?” usually asks about an average

Many intervention studies estimate an average effect across the participants included in the analysis. That estimate is useful. It tells researchers whether outcomes differ, on average, between relevant conditions under the assumptions of the design and analysis.

But an average does not imply that everyone experiences the average effect. Effects may vary across individuals, populations, settings, intervention characteristics, doses, durations, baseline risks, or other conditions.

This variation is often described as treatment-effect heterogeneity or, more generally, effect heterogeneity. A variable associated with differences in the effect may be described as an effect modifier or moderator, depending on the methodological framework being used.

Average effect The estimated effect summarized across the population represented by the analysis.
Effect heterogeneity Variation in the magnitude, and sometimes direction, of effects across people, groups, settings, or other conditions.

Move toward heterogeneity when the existence of an average effect is no longer the main uncertainty

The transition becomes especially relevant when repeated credible studies or evidence syntheses support a reasonably stable overall conclusion. At that point, another study designed primarily to demonstrate the same average effect in essentially the same circumstances may provide diminishing information.

This does not mean the average effect becomes irrelevant. Rather, the research frontier may move. If new studies rarely change the overall conclusion, understanding meaningful variation can become more informative than repeatedly establishing the same central tendency.

“For whom?” asks whether effects differ across relevant people or groups

An intervention may not benefit all participants equally. Baseline characteristics, prior risk, age, disease severity, prior knowledge, socioeconomic conditions, or other features may be associated with different effects, depending on the field and question.

Researchers should be cautious here. Cochrane guidance emphasizes that subgroup analyses have substantial pitfalls. Conducting many post hoc subgroup comparisons can generate misleading apparent differences simply by chance. Comparing whether one subgroup has a statistically significant effect while another does not is also not a valid test that the effects differ between the groups.

Credible claims about differential effects therefore require appropriate interaction analyses, adequate information, preferably prespecified hypotheses, substantive plausibility, and careful interpretation.

“When?” asks about conditions and boundary effects

Variation does not have to be a property of people. Effects can depend on intervention intensity, duration, timing, setting, delivery mode, organizational conditions, comparator, follow-up period, or other contextual characteristics.

Identifying such boundary conditions can turn a broad statement such as “the intervention works” into a more useful statement about the circumstances in which a particular effect should be expected.

This can also explain apparent inconsistency in an accumulating literature. Studies that seem to disagree may actually have investigated meaningfully different conditions.

“Why?” can refer to mechanisms, but moderators and mechanisms are not the same

A moderator identifies variation in an effect. A mechanism concerns the process through which an exposure or intervention produces an outcome. These questions can be related, but they should not be collapsed into one another.

Suppose an educational intervention has a larger effect among students with lower baseline knowledge. Baseline knowledge may help identify for whom the intervention is particularly beneficial, but that observation alone does not explain the causal process through which the intervention improves learning.

When an association itself has become well established, researchers may need a more focused transition from association toward mechanism.

Question What it asks Example
Does it work? Is there an average effect under the studied conditions? Does the intervention improve learning outcomes on average?
For whom? Does the effect differ across relevant people or groups? Does prior knowledge modify the effect?
When? Under which conditions does the effect become larger, smaller, or absent? Does delivery intensity alter the effect?
Why? What process or mechanism could account for the effect? Does the intervention improve learning by increasing retrieval practice?

Heterogeneity should be explained carefully, not mined until something appears

Once researchers become interested in “for whom” and “when,” it is tempting to divide data into many subgroups and search for differences. That strategy can produce convincing-looking but unreliable findings.

Cochrane recommends caution when investigating heterogeneity through subgroup analyses and meta-regression. Analyses specified after inspecting the results are particularly vulnerable to spurious explanations and are better treated as hypothesis-generating. Even prespecified analyses require careful interpretation.

The shift toward heterogeneity therefore demands stronger questions, not simply more analyses.

Watch Out

Do not infer that effects differ between two groups merely because the effect is statistically significant in one group and not statistically significant in the other. The difference between groups needs to be examined directly.

A null average effect can also hide important heterogeneity

The move toward heterogeneity is not restricted to interventions with a positive average effect. An overall estimate near zero can sometimes arise because beneficial and harmful effects, or larger and smaller effects, occur under different conditions.

That possibility should not be used as an excuse to search indiscriminately for subgroups after an unpromising result. It becomes scientifically useful when there are credible reasons, suitable data, and appropriate methods for expecting differential effects.

The shift does not require waiting until everything about the average effect is certain

Research development is rarely perfectly sequential. Questions about heterogeneity and mechanisms can be planned while effectiveness evidence is still accumulating. Indeed, waiting until a long sequence of conventional studies is complete may waste opportunities to learn about context and variation.

The relevant issue is whether there is enough evidence and theoretical justification to support the more specific question. Research on effect modification requires adequate variation, measurement, sample size, and design. Asking a sophisticated question with data incapable of answering it merely upgrades the vocabulary, not the evidence.

The next transition may be toward real-world implementation

Once researchers understand that an intervention can produce benefits and have begun identifying important conditions surrounding those benefits, another question may become pressing: can the intervention be adopted, delivered, and sustained in routine settings?

That requires a different shift toward implementation research. Knowing for whom something works does not automatically tell you how to make it work in practice.

04 · A Practical Example

Moving beyond the average effect of a learning intervention

Hypothetical Example

An intervention with a repeatedly observed average benefit

Imagine that several reasonably rigorous studies find that a structured learning intervention improves examination performance compared with usual instruction. A synthesis of the studies supports a modest average benefit, and additional studies have not substantially altered that broad conclusion.

Original question Does the intervention improve examination performance on average?
Accumulated answer The evidence reasonably supports an average benefit under the conditions studied.
Emerging uncertainty Effects appear to vary across students and instructional conditions, but the reasons for that variation are unclear.
Next research questions Do effects differ by baseline knowledge? Does intervention intensity matter? Which learning processes mediate improvement? Are apparent subgroup differences reproducible?

Another conventional effectiveness study could still contribute information. But if it simply reproduces the same average comparison in a familiar population, it may answer a question that the literature already addresses reasonably well while leaving the more consequential uncertainty untouched.

05 · What Researchers Often Get Wrong

Moving beyond “Does it work?” requires more than adding subgroup analyses

Misconception

If the average effect is significant, researchers should immediately study subgroups

Statistical significance does not establish that a literature is ready for reliable heterogeneity analysis. Researchers also need substantive hypotheses, suitable variation, adequate information, appropriate measurements, and methods capable of testing differences in effects.

Misconception

If one subgroup has a significant effect and another does not, the effect differs between them

This comparison is invalid. The relevant question is whether the effect estimates themselves differ, which requires an appropriate test of interaction or another suitable comparison.

Misconception

Finding a moderator explains why an intervention works

A moderator identifies conditions associated with different effects. It does not necessarily identify the causal mechanism producing the effect. Mechanism requires a different inferential argument.

Misconception

An average effect applies equally to everyone

An average summarizes a distribution of effects under the relevant model and population. Individual or subgroup effects may differ from that average, sometimes substantially.

Misconception

Every subgroup difference deserves a separate study

Apparent differences can arise by chance, especially when many subgroups are examined. Prioritize differences that are theoretically plausible, consequential, appropriately tested, and supported by evidence rather than those discovered through unrestricted searching.

06 · What This Means for You

Let the remaining uncertainty determine the next question

If the literature already provides credible evidence about an average effect, examine what researchers still cannot explain. Do outcomes vary meaningfully? Are important populations poorly represented? Are plausible effect modifiers known but inadequately tested? Is the mechanism uncertain?

A simple decision framework

If the existence or direction of the average effect remains genuinely uncertain
Further rigorous effectiveness research may still deserve priority.
If the average effect is reasonably established but meaningful variation remains unexplained
Design research specifically capable of testing credible effect modifiers and boundary conditions.
If moderators are increasingly well characterized but the causal process remains unclear
Consider designs that directly test competing mechanisms rather than treating moderation as explanation.
If effectiveness is credible and the main uncertainty concerns routine delivery
Implementation questions may now have greater practical value.

The broader principle is simple: research questions should evolve with accumulated knowledge. Once a literature has repeatedly answered one question, novelty should not depend on asking it again with a slightly different sample. The more useful contribution may be identifying what becomes important once the original question is largely answered.

07 · A Quick Checklist

Decide whether the literature is ready for more specific questions

Before moving beyond “Does it work?”, check:
Verify whether credible accumulated evidence already supports a reasonably stable conclusion about the average effect.
Identify meaningful variation that the average effect may conceal.
Specify potential moderators or boundary conditions using substantive reasoning rather than unrestricted post hoc searching.
Ensure the study contains enough information and variation to test differential effects credibly.
Use appropriate interaction analyses rather than comparing significance within separate subgroups.
Distinguish moderators from mechanisms and design the study according to the question actually being asked.
Consider whether implementation rather than another effectiveness study has become the more consequential uncertainty.
08 · Frequently Asked Questions

Questions about moving beyond average effects

Do researchers need to prove an average effect before studying moderators?

No universal sequence is required. Heterogeneity can sometimes be scientifically important even when the average effect is small or uncertain. The moderator question should nevertheless be theoretically justified and supported by a design capable of answering it.

What is treatment-effect heterogeneity?

It refers to variation in treatment or intervention effects across individuals, groups, or conditions. The same intervention may therefore have effects of different magnitudes, or occasionally different directions, across relevant circumstances.

Is a moderator the same as a mediator?

No. A moderator concerns whether or under what conditions an effect differs. A mediator concerns a process through which an effect may occur. The statistical and causal assumptions involved also differ.

Can I identify moderators by running many subgroup analyses?

You can explore potential differences, but unrestricted subgroup searching carries a substantial risk of false findings. Prespecified, theoretically plausible analyses with appropriate tests and adequate information provide stronger evidence.

What if the average effect is close to zero?

A near-zero average does not prove that effects are identical for everyone. Meaningful heterogeneity is possible, but investigating it requires a substantive rationale and appropriate evidence rather than searching for a subgroup in which the result becomes favorable.

Does studying “for whom” mean research has become personalized?

Not automatically. Detecting credible variation in effects is one step. Making reliable predictions for particular individuals generally requires stronger evidence, validation, and methods suited to individualized inference.

09 · The Bottom Line

Once the average answer is credible, variation may become the more important science

The Bottom Line

Research should increasingly move from “Does it work?” to “For whom, when, and why?” when the average effect is reasonably established and meaningful uncertainty has shifted toward effect heterogeneity, boundary conditions, or explanation.

That transition should produce more precise hypotheses, not a fishing expedition through subgroups. The aim is to understand meaningful variation that the average obscures and to ask questions that accumulated evidence has made both possible and necessary.

10 · Sources and Further Reading

Sources and further reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes