Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How Do You Recognize When Authors Treat Statistical Significance as Practical Importance?

A statistically significant result is not automatically large, useful, or important. Learn how to recognize when authors use a p-value as evidence of practical importance and what to examine instead.

332
Statistical vs. Practical Importance Guide 332 of 899
01 · The Question

Does a statistically significant result actually matter?

A study reports p <.001. The authors call the result important. The abstract describes a meaningful improvement. The Discussion recommends changing practice.

Those conclusions may ultimately be justified, but the p-value did not establish them.

Statistical significance addresses a statistical question about the compatibility of the observed data with a specified model. Practical importance asks a different question: is the estimated effect large or consequential enough to matter for the scientific, educational, clinical, organizational, policy, or other decision being considered?

Confusing those questions is one of the most persistent ways a technically correct statistical result can acquire more substantive importance than the evidence actually demonstrates.

02 · The Short Answer

Statistical significance does not establish practical importance

In Brief

You can recognize this mistake when authors use statistical significance, often a p-value below a conventional threshold, as the main reason for describing a result as important, meaningful, substantial, useful, or consequential without separately evaluating the magnitude of the effect and its relevance to the research context.

To judge practical importance, examine the effect estimate, its uncertainty, the measurement scale, a substantively justified benchmark where available, baseline risk or absolute difference when relevant, potential benefits and harms, and the decision the result is supposed to inform.

03 · What You Need to Know

Why statistical significance and practical importance answer different questions

A p-value does not tell you how large an effect is

The American Statistical Association explicitly states that a p-value or statistical significance does not measure the size of an effect or the importance of a result. It also cautions against basing scientific conclusions or decisions only on whether a p-value crosses a particular threshold.

This means that p <.05 cannot, by itself, tell you whether a difference is educationally meaningful, clinically consequential, economically worthwhile, socially important, or useful in practice.

The p-value and the substantive magnitude answer different questions.

Statistical significance Concerns the result of a statistical procedure under a specified model and threshold; it does not measure the magnitude or practical importance of the effect.
Practical importance Concerns whether the magnitude and consequences of an effect matter for the substantive question, population, or decision at hand.

Large samples can make small differences statistically detectable

Statistical detectability depends partly on the amount of information available. With sufficiently large samples and sufficiently precise measurements, relatively small differences can produce very small p-values.

Suppose an educational intervention raises the mean score on a 100-point assessment from 78.40 to 78.75. With a very large sample, that 0.35-point difference might be estimated precisely enough to meet a conventional significance threshold.

The statistical result would still leave a substantive question unanswered: does a 0.35-point difference matter?

The answer depends on the assessment, its measurement properties, the educational context, the costs and burdens of the intervention, and what magnitude would alter a meaningful outcome or decision.

Start with the effect estimate

When authors describe a significant result as important, locate the actual effect.

Depending on the analysis, this might be a mean difference, standardized mean difference, risk difference, risk ratio, odds ratio, hazard ratio, correlation coefficient, regression coefficient, change in probability, or another quantity.

Then ask what that quantity means on the substantive scale.

A regression coefficient of 0.18 can look small or large depending on the variables and units involved. A two-point difference means little until you know whether the scale runs from 0 to 5, 0 to 100, or 0 to 1,000 and what those points represent.

This is why separating the numerical result from the authors’ narrative is useful before judging importance.

Effect size is necessary information, but it does not interpret itself

Effect sizes help quantify magnitude, but attaching a standardized effect-size label does not automatically settle practical importance.

Generic descriptors such as “small,” “medium,” or “large” can sometimes provide a rough reference, but their substantive meaning depends on context. A statistically small effect can matter when an intervention is inexpensive, scalable, cumulative, or applied to a consequential outcome across many people. A statistically large effect can be less useful when the outcome is weakly relevant, the estimate is highly uncertain, or implementation is impractical.

The better question is not merely “Is the effect size large?” but “Is this magnitude consequential for this outcome and this decision?”

Confidence intervals help you judge which magnitudes remain plausible

A point estimate alone does not show how precisely an effect has been estimated.

Suppose an intervention produces an estimated improvement of 4 points. A 95% confidence interval from 3.5 to 4.5 suggests a substantially different degree of precision from an interval ranging from -1 to 9 points.

If practical importance begins around a substantively justified 5-point difference, those intervals also have different implications. The first is relatively precise but centered below that benchmark. The second remains compatible with effects ranging from negligible or unfavorable to potentially important.

CONSORT 2025 emphasizes reporting effect estimates with confidence intervals and explicitly cautions that statistical significance and clinical importance should not be confused.

Practical importance requires a substantive benchmark

Whenever possible, ask what magnitude would actually matter before looking at whether the observed result crossed a significance threshold.

Different fields use different criteria. Clinical research may use a minimally important difference or clinically relevant risk reduction. Education might consider whether an achievement difference corresponds to a meaningful change in mastery, progression, failure rates, or another consequential outcome. Organizational research might consider productivity, retention, costs, or decision thresholds.

Not every field has a well-established benchmark, and apparently precise thresholds may themselves be contested. When no accepted threshold exists, practical importance requires explicit substantive reasoning rather than a hidden substitution of p <.05 for “important.”

Relative effects can sound more important than absolute effects

Consider an outcome whose probability falls from 2% to 1%.

That can be described as a 50% relative reduction or a 1-percentage-point absolute reduction. Both descriptions can be mathematically correct, yet they foreground different aspects of the result.

Which representation matters more depends on the decision context. If the outcome is severe and the intervention inexpensive and safe, a one-percentage-point reduction could be important. If the intervention is costly or harmful, the same absolute difference might lead to a different judgment.

For binary outcomes in randomized trials, CONSORT 2025 recommends reporting both absolute and relative effect sizes where appropriate because each contributes useful information.

Importance depends partly on consequences, not only magnitude

Practical importance is not encoded entirely inside an effect estimate.

Imagine two interventions producing the same average improvement. One is inexpensive, requires five minutes, and has negligible adverse consequences. The other requires substantial funding, specialized staff, and considerable participant burden.

The same effect magnitude can support different decisions.

This is why claims such as “the statistically significant improvement supports widespread implementation” often require evidence beyond statistical significance itself. Costs, harms, feasibility, equity, alternatives, and durability may all matter.

A non-significant result can still be practically important

The confusion also works in reverse.

Suppose an intervention has an estimated benefit large enough to matter, but its confidence interval is wide because the study is small. The result may not meet a conventional significance threshold, yet potentially important effects remain compatible with the data.

That does not mean the intervention has been shown to work. It means “not statistically significant” should not be translated into “unimportant.” This is closely related to the mistake of turning insufficient evidence into evidence of no effect.

Watch for importance being added in the Discussion

The Results may simply report p =.003 and an effect estimate. The Discussion then describes the result as “substantial,” “meaningful,” or “highly important.”

Those adjectives require a separate basis.

Ask what criterion makes the effect substantial. Does the paper cite an established threshold? Compare the magnitude with prior evidence? Explain consequences for participants or decisions? Consider absolute effects? Or does the interpretation rest primarily on statistical significance?

If the last explanation fits best, the Discussion may be making the result sound stronger than its substantive magnitude warrants.

Watch Out

Do not assume that every numerically small effect is practically trivial. Small effects can accumulate, affect large populations, alter consequential outcomes, or be worthwhile when interventions are low-cost and low-risk. Practical importance must be argued in context rather than inferred from either the p-value or the apparent size of a number alone.

04 · A Practical Example

A tiny p-value can accompany a tiny difference

Hypothetical Example

A large educational study with a very precise result

Imagine a randomized study involving 20,000 students. The intervention is intended to improve performance on a 100-point examination.

Observed result Mean score: 80.35 in the intervention group and 80.00 in the comparison group.
Estimated effect Mean difference: 0.35 points, with a narrow confidence interval excluding zero.
Statistical result The corresponding hypothesis test produces p <.001.
Overextended interpretation “The intervention produced a highly significant and important improvement in student achievement.”
Question still unanswered Is a 0.35-point improvement on this assessment large enough to matter educationally, and is it worth the intervention’s cost, effort, and potential trade-offs?

The study may provide strong statistical evidence that the average difference is not zero under the specified analysis. That does not automatically establish that the difference is educationally consequential.

If 0.35 points changes an important pass-fail decision for many students, the result might matter considerably. If it produces no consequential difference and requires an expensive intervention, the practical interpretation could be quite different. The p-value cannot choose between those possibilities.

05 · What Researchers Often Get Wrong

Common mistakes when judging statistical and practical importance

Misconception

A smaller p-value means a larger effect

No. The p-value depends on the observed effect in relation to its uncertainty under the statistical model. Sample size and variability matter, so a small effect estimated very precisely can have a smaller p-value than a much larger but imprecisely estimated effect.

Misconception

p <.001 means the result is extremely important

The number indicates neither practical importance nor effect magnitude. Importance must be evaluated using the substantive meaning and consequences of the estimated effect.

Misconception

A large standardized effect is automatically important

Not necessarily. Standardized effect sizes need substantive interpretation. Outcome relevance, uncertainty, costs, harms, baseline conditions, and the decision context can all affect whether an estimated magnitude matters.

Misconception

A non-significant effect is practically unimportant

No. An imprecise study may fail to reach a significance threshold while remaining compatible with effects that would matter considerably. Examine the estimate and uncertainty.

Misconception

Practical significance is simply another statistical test

Practical importance is fundamentally substantive. Statistical tools can help quantify whether effects cross meaningful thresholds, but the threshold itself must be justified by the research context rather than generated automatically by a p-value.

06 · What This Means for You

Ask “How much?” before asking “Was it significant?”

When a paper emphasizes statistical significance, temporarily ignore the adjective “significant” and reconstruct the substantive result.

A simple importance check

If p <.05
Find the effect estimate and ask whether its magnitude matters on the relevant substantive scale.
If the effect is described as meaningful
Identify the benchmark, consequence, or decision criterion that makes it meaningful.
If the estimate is uncertain
Use the confidence interval or other uncertainty measure to examine which practically relevant magnitudes remain compatible with the evidence.
If relative effects sound dramatic
Inspect absolute risks, rates, counts, or differences where relevant.
If implementation is recommended
Consider whether benefits, harms, costs, feasibility, alternatives, and other decision-relevant factors support that recommendation.

Then return to the distinction between what the data show and what the authors conclude. Statistical evidence can support an effect. The judgment that the effect matters requires another layer of reasoning.

07 · A Quick Checklist

Does the result matter beyond its p-value?

Before calling a statistically significant result important, check:
Identify the actual effect estimate rather than relying on the p-value.
Interpret the magnitude on the original or otherwise substantively meaningful scale where possible.
Examine the confidence interval or other representation of uncertainty.
Identify a defensible benchmark for what would count as a meaningful difference, if one exists.
Check absolute as well as relative effects when both are relevant.
Consider whether the outcome itself is consequential for the research question.
Consider costs, harms, burden, feasibility, and alternatives when practical decisions are involved.
Check whether words such as “important,” “meaningful,” or “substantial” have justification beyond statistical significance.
08 · Frequently Asked Questions

Questions about statistical significance and practical importance

Does p <.05 mean the effect is important?

No. A significance threshold does not measure effect magnitude or practical importance. Examine the estimated effect, uncertainty, substantive scale, and consequences.

Does a smaller p-value mean a bigger effect?

No. P-values are influenced by effect magnitude, variability, sample size, and the statistical model. Very small effects can produce very small p-values when estimated precisely.

Is practical significance the same as effect size?

No. Effect size quantifies magnitude. Practical importance is a judgment about whether that magnitude matters for the substantive question or decision. Effect size is therefore important evidence for practical significance, but the two are not synonyms.

Can a small effect be practically important?

Yes. Small effects can matter when outcomes are consequential, exposure is widespread, effects accumulate, or interventions are inexpensive and low-risk. Context determines whether the magnitude is consequential.

Can a statistically non-significant result still be practically important?

Potentially. The point estimate may be large enough to matter while uncertainty remains too great for a conventional significance threshold. That is not proof of an important effect, but it means practical importance cannot be dismissed from the p-value alone.

What if there is no established threshold for a meaningful effect?

Then practical interpretation should be explicit about the absence of a universally accepted threshold. Consider the measurement scale, prior evidence, stakeholder consequences, costs, harms, and plausible decision thresholds rather than inventing a precise cutoff.

09 · The Bottom Line

Statistically detectable does not automatically mean consequential

The Bottom Line

Authors treat statistical significance as practical importance when they use a p-value or significance threshold as the main basis for calling an effect meaningful, important, substantial, or useful without separately evaluating its magnitude and substantive consequences.

Read the effect estimate, uncertainty, scale, and relevant practical benchmark, then consider the decision context. Statistical significance can contribute to interpretation, but it cannot decide for you whether an effect actually matters.

10 · Sources and Further Reading

Sources and further reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes