01 · The Question
When Is Looking at Subgroups Good Science, and When Is It Fishing?
An overall analysis shows little evidence of an effect. You divide the sample by age and find something interesting among younger participants. You then examine men and women, high- and low-performing participants, different regions, several baseline categories, and perhaps combinations of them.
Eventually, one subgroup produces a striking result.
Subgroup analysis can answer important scientific questions about whether an association or intervention effect differs across populations. But as the number of possible groups increases, so does the opportunity to discover an apparently compelling pattern by chance. The critical question is whether subgroup analyses are being used to investigate credible heterogeneity or searched selectively until a favorable result appears.
03 · What You Need to Know
The Statistical Question Is Whether the Groups Differ From Each Other
A Subgroup Effect Means the Effect Differs Across Groups
Researchers conduct subgroup analyses when they want to know whether an effect varies according to a characteristic such as age, sex, baseline severity, prior exposure, educational level, location, or another scientifically relevant factor.
In an intervention study, for example, the substantive question might be whether the intervention works differently for younger and older participants.
That question is not answered simply by testing the intervention separately within each group.
The relevant statistical question is whether the estimated intervention effects differ from each other. In many settings, this is evaluated using an interaction between subgroup membership and the exposure or treatment. CONSORT 2025 specifically describes subgroup comparisons in randomized trials in terms of tests of interaction and recommends reporting estimated differences in intervention effects with confidence intervals when formal interaction analyses are undertaken.
“Significant Here but Not There” Does Not Establish a Subgroup Difference
This is one of the most consequential mistakes in subgroup analysis.
Suppose an intervention produces:
-
Younger participants: estimated effect = 4.2 points, p =.03
-
Older participants: estimated effect = 3.8 points, p =.08
It is tempting to conclude that the intervention works for younger participants but not older participants.
That conclusion does not follow from the two p-values. The estimated effects, 4.2 and 3.8, are actually quite similar. Different p-values can arise because the subgroups have different sample sizes, variability, or precision.
As Altman and Bland famously explained, the appropriate comparison concerns the difference between the estimates. CONSORT likewise warns that one statistically significant subgroup result and one nonsignificant subgroup result do not establish an interaction.
Watch Out
“Significant in group A but not significant in group B” is not equivalent to “the effect is significantly different between groups A and B.” If your claim concerns a subgroup difference, analyze the difference directly.
Searching More Subgroups Creates More Opportunities for Chance Findings
Suppose you examine whether an effect differs by sex, four age categories, three educational levels, five regions, baseline achievement, prior experience, socioeconomic status, and several combinations of those variables.
Even if no genuine subgroup effect exists, repeated analyses provide opportunities for apparently unusual results to emerge through sampling variation.
This is part of the broader multiple-comparisons problem. CONSORT 2025 cautions that multiple analyses of the same data create risks of false-positive findings and specifically advises researchers to resist performing numerous subgroup analyses.
The problem becomes more severe when only the subgroup producing the most favorable result appears in the manuscript. Readers then see one apparent discovery without seeing the analytical search from which it was selected.
Prespecified Subgroups Generally Carry More Credibility
A subgroup hypothesis specified before examining the relevant results has a different evidential history from one discovered after searching the dataset.
Suppose previous research provides a strong reason to expect an intervention to work differently according to baseline severity. You specify that interaction in the protocol and analysis plan before examining the outcome data.
Now compare that with dividing participants repeatedly by whatever characteristics happen to produce interesting differences after the main analysis is complete.
The first analysis has a prospectively stated rationale. The second is exploratory.
CONSORT recommends reporting which subgroup analyses were conducted, why they were undertaken, whether they were prespecified, and how many were prespecified. It notes that analyses suggested by the data generally carry less credibility than those planned in advance.
Prespecification does not make a subgroup effect automatically true. Poorly justified subgroup hypotheses can be preregistered too. It simply removes one important route by which the observed results themselves can determine which subgroup receives attention.
Post Hoc Subgroup Analysis Is Not Automatically Forbidden
An unexpected overall finding may raise a scientifically interesting question about heterogeneity. Exploring that possibility can be worthwhile.
The problem is not discovery. It is pretending discovery was confirmation.
If an unexpected pattern leads you to examine a subgroup after seeing the data, report the analysis as post hoc or exploratory where that distinction affects interpretation. Explain how the subgroup was chosen and avoid presenting the finding with the same evidential weight as a strong prespecified test.
CONSORT 2025 notes that post hoc subgroup comparisons are especially unlikely to be confirmed and that many subgroup claims lack strong statistical support or independent corroboration.
This is another application of the broader principle that researchers can report exploratory findings without pretending they were predicted.
Data-Driven Cutoffs Create Another Layer of Flexibility
Continuous variables such as age, income, baseline scores, or symptom severity can be divided into groups in many ways.
Age could become under 30 versus 30 and older, under 40 versus 40 and older, three age bands, quartiles, or dozens of other classifications.
If researchers try several cutoffs and retain whichever produces the strongest interaction, the subgroup definition itself has become part of the analytical search.
CONSORT warns against choosing cut points according to statistical significance. It also notes that categorizing continuous variables can discard information and reduce statistical power, particularly when arbitrary cutoffs lack substantive justification.
When feasible, modeling how an effect changes across the continuous variable may sometimes be preferable to inventing an arbitrary boundary solely to create groups.
A Plausible Explanation Afterward Does Not Make the Subgroup Prespecified
Suppose you unexpectedly find an effect among participants younger than 30. After seeing the result, you construct a convincing theoretical explanation involving developmental differences.
The explanation may be genuinely insightful. It does not change when the hypothesis was generated.
Presenting the theory as though it motivated the subgroup analysis beforehand would introduce a HARKing problem. The more defensible approach is to say that the subgroup pattern was unexpected, explain the possible mechanism, and identify independent confirmation as an important next step.
Subgroup Claims Should Be Judged on More Than One P-Value
The credibility of an apparent subgroup effect depends on the broader evidence.
Useful considerations include whether the subgroup variable was specified in advance, whether there was a strong rationale, whether the interaction itself is supported, how many subgroup analyses were performed, whether the direction was predicted, whether the effect is consistent with other evidence, and whether it has been independently replicated.
Sun and colleagues developed criteria for assessing the credibility of subgroup effects precisely because statistically significant subgroup findings can be spurious and yet potentially influential for clinical or policy decisions.
No single criterion guarantees that an apparent interaction is genuine. Subgroup findings should generally be interpreted as stronger when several credibility indicators align rather than because one analysis happens to cross p =.05.
04 · A Practical Example
How an Overall Null Result Can Produce a Tempting Subgroup Story
Hypothetical Example
An Educational Intervention With No Clear Overall Effect
A randomized study tests whether a digital tutoring program improves examination performance. The overall treatment effect is small and statistically nonsignificant.
First subgroup
The researchers compare male and female students. Neither interaction nor pattern is particularly striking.
More subgroups
They examine year level, prior achievement, socioeconomic status, campus, program, and several age cutoffs.
A promising result appears
Among students younger than 20, the intervention effect produces p =.02.
The tempting conclusion
The researchers report that the intervention “works for younger students” while giving little indication that numerous subgroup definitions were searched.
That conclusion has several problems.
First, the researchers need to test whether the estimated effect among younger students actually differs from the effect among older students. A significant result in one subgroup and a nonsignificant result in another would not establish that difference.
Second, the subgroup emerged from a broad post hoc search. Its p-value cannot be interpreted as though “students younger than 20” had been the single subgroup hypothesis selected independently of the data.
The result can still be reported. A more appropriate conclusion might identify it as an exploratory age-related pattern that requires confirmation in independent data rather than evidence that the intervention has already been shown to work specifically for younger students.