01 · The Question
Does a statistically significant result actually matter?
A study reports p <.001. The authors call the result important. The abstract describes a meaningful improvement. The Discussion recommends changing practice.
Those conclusions may ultimately be justified, but the p-value did not establish them.
Statistical significance addresses a statistical question about the compatibility of the observed data with a specified model. Practical importance asks a different question: is the estimated effect large or consequential enough to matter for the scientific, educational, clinical, organizational, policy, or other decision being considered?
Confusing those questions is one of the most persistent ways a technically correct statistical result can acquire more substantive importance than the evidence actually demonstrates.
03 · What You Need to Know
Why statistical significance and practical importance answer different questions
A p-value does not tell you how large an effect is
The American Statistical Association explicitly states that a p-value or statistical significance does not measure the size of an effect or the importance of a result. It also cautions against basing scientific conclusions or decisions only on whether a p-value crosses a particular threshold.
This means that p <.05 cannot, by itself, tell you whether a difference is educationally meaningful, clinically consequential, economically worthwhile, socially important, or useful in practice.
The p-value and the substantive magnitude answer different questions.
Statistical significance
Concerns the result of a statistical procedure under a specified model and threshold; it does not measure the magnitude or practical importance of the effect.
Practical importance
Concerns whether the magnitude and consequences of an effect matter for the substantive question, population, or decision at hand.
Large samples can make small differences statistically detectable
Statistical detectability depends partly on the amount of information available. With sufficiently large samples and sufficiently precise measurements, relatively small differences can produce very small p-values.
Suppose an educational intervention raises the mean score on a 100-point assessment from 78.40 to 78.75. With a very large sample, that 0.35-point difference might be estimated precisely enough to meet a conventional significance threshold.
The statistical result would still leave a substantive question unanswered: does a 0.35-point difference matter?
The answer depends on the assessment, its measurement properties, the educational context, the costs and burdens of the intervention, and what magnitude would alter a meaningful outcome or decision.
Start with the effect estimate
When authors describe a significant result as important, locate the actual effect.
Depending on the analysis, this might be a mean difference, standardized mean difference, risk difference, risk ratio, odds ratio, hazard ratio, correlation coefficient, regression coefficient, change in probability, or another quantity.
Then ask what that quantity means on the substantive scale.
A regression coefficient of 0.18 can look small or large depending on the variables and units involved. A two-point difference means little until you know whether the scale runs from 0 to 5, 0 to 100, or 0 to 1,000 and what those points represent.
This is why separating the numerical result from the authors’ narrative is useful before judging importance.
Effect size is necessary information, but it does not interpret itself
Effect sizes help quantify magnitude, but attaching a standardized effect-size label does not automatically settle practical importance.
Generic descriptors such as “small,” “medium,” or “large” can sometimes provide a rough reference, but their substantive meaning depends on context. A statistically small effect can matter when an intervention is inexpensive, scalable, cumulative, or applied to a consequential outcome across many people. A statistically large effect can be less useful when the outcome is weakly relevant, the estimate is highly uncertain, or implementation is impractical.
The better question is not merely “Is the effect size large?” but “Is this magnitude consequential for this outcome and this decision?”
Confidence intervals help you judge which magnitudes remain plausible
A point estimate alone does not show how precisely an effect has been estimated.
Suppose an intervention produces an estimated improvement of 4 points. A 95% confidence interval from 3.5 to 4.5 suggests a substantially different degree of precision from an interval ranging from -1 to 9 points.
If practical importance begins around a substantively justified 5-point difference, those intervals also have different implications. The first is relatively precise but centered below that benchmark. The second remains compatible with effects ranging from negligible or unfavorable to potentially important.
CONSORT 2025 emphasizes reporting effect estimates with confidence intervals and explicitly cautions that statistical significance and clinical importance should not be confused.
Practical importance requires a substantive benchmark
Whenever possible, ask what magnitude would actually matter before looking at whether the observed result crossed a significance threshold.
Different fields use different criteria. Clinical research may use a minimally important difference or clinically relevant risk reduction. Education might consider whether an achievement difference corresponds to a meaningful change in mastery, progression, failure rates, or another consequential outcome. Organizational research might consider productivity, retention, costs, or decision thresholds.
Not every field has a well-established benchmark, and apparently precise thresholds may themselves be contested. When no accepted threshold exists, practical importance requires explicit substantive reasoning rather than a hidden substitution of p <.05 for “important.”
Relative effects can sound more important than absolute effects
Consider an outcome whose probability falls from 2% to 1%.
That can be described as a 50% relative reduction or a 1-percentage-point absolute reduction. Both descriptions can be mathematically correct, yet they foreground different aspects of the result.
Which representation matters more depends on the decision context. If the outcome is severe and the intervention inexpensive and safe, a one-percentage-point reduction could be important. If the intervention is costly or harmful, the same absolute difference might lead to a different judgment.
For binary outcomes in randomized trials, CONSORT 2025 recommends reporting both absolute and relative effect sizes where appropriate because each contributes useful information.
Importance depends partly on consequences, not only magnitude
Practical importance is not encoded entirely inside an effect estimate.
Imagine two interventions producing the same average improvement. One is inexpensive, requires five minutes, and has negligible adverse consequences. The other requires substantial funding, specialized staff, and considerable participant burden.
The same effect magnitude can support different decisions.
This is why claims such as “the statistically significant improvement supports widespread implementation” often require evidence beyond statistical significance itself. Costs, harms, feasibility, equity, alternatives, and durability may all matter.
A non-significant result can still be practically important
The confusion also works in reverse.
Suppose an intervention has an estimated benefit large enough to matter, but its confidence interval is wide because the study is small. The result may not meet a conventional significance threshold, yet potentially important effects remain compatible with the data.
That does not mean the intervention has been shown to work. It means “not statistically significant” should not be translated into “unimportant.” This is closely related to the mistake of turning insufficient evidence into evidence of no effect.
Watch for importance being added in the Discussion
The Results may simply report p =.003 and an effect estimate. The Discussion then describes the result as “substantial,” “meaningful,” or “highly important.”
Those adjectives require a separate basis.
Ask what criterion makes the effect substantial. Does the paper cite an established threshold? Compare the magnitude with prior evidence? Explain consequences for participants or decisions? Consider absolute effects? Or does the interpretation rest primarily on statistical significance?
If the last explanation fits best, the Discussion may be making the result sound stronger than its substantive magnitude warrants.
Watch Out
Do not assume that every numerically small effect is practically trivial. Small effects can accumulate, affect large populations, alter consequential outcomes, or be worthwhile when interventions are low-cost and low-risk. Practical importance must be argued in context rather than inferred from either the p-value or the apparent size of a number alone.