Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Questions Remain Unanswered Because Research Has Focused on Statistical Significance Rather Than Meaningful Effects?

A literature can contain many statistically significant findings while leaving the important question unanswered: how large is the effect, how precisely is it known, and does it actually matter?

652
Beyond Statistical Significance Guide 652 of 899
01 · The Question

What If the Literature Keeps Asking Whether an Effect Exists Instead of Whether It Matters?

You review a mature literature and encounter the same language repeatedly: significant, non-significant, significant, significant again.

At first, this can make the evidence appear decisive. If enough studies report p <.05, surely the important question has been answered.

Not necessarily.

A significance test can contribute information about how compatible observed data are with a specified statistical model or null hypothesis. It does not, by itself, tell you whether an effect is large, useful, consequential, precisely estimated, theoretically important, or worth acting on. The American Statistical Association has explicitly cautioned against treating a p-value or a statistical-significance threshold as a measure of the size or importance of an effect.

A literature organized primarily around whether effects cross a significance threshold can therefore accumulate publications while leaving a more consequential question unresolved: how much difference does the phenomenon actually make?

02 · The Short Answer

Move From “Is There an Effect?” to “How Large and Meaningful Is It?”

In Brief

Questions remain unanswered when research focuses on statistical significance rather than meaningful effects because p-value thresholds do not establish the magnitude, precision, practical importance, or consequences of an observed effect.

A stronger evidence base reports and interprets effect estimates together with their uncertainty and evaluates them against substantively meaningful differences. The relevant threshold for “meaningful” depends on the outcome, research domain, decision, stakeholders, benefits, harms, costs, and context.

03 · What You Need to Know

Statistical Detectability and Substantive Importance Are Different Questions

A p-value does not tell you how large an effect is

The American Statistical Association's statement on p-values emphasizes that statistical significance does not measure the size of an effect or the importance of a result. This distinction is foundational.

Suppose an educational intervention improves examination performance by 0.2 percentage points in a very large sample and the result produces a very small p-value. The data may provide evidence against a precise null hypothesis of no difference. That does not tell you whether a 0.2-point improvement matters educationally.

Conversely, a smaller study could estimate an improvement large enough to matter but with substantial uncertainty, producing a p-value above.05. Labeling the first result a “success” and the second “no effect” would obscure what the studies actually estimated.

Statistical significance A threshold-based interpretation of a statistical test, commonly based on whether a p-value falls below a conventional cutoff.
Meaningful effect An effect large enough to matter for the substantive scientific, practical, clinical, educational, policy, or other decision under consideration.

Sample size can change significance without changing the substantive effect

P-values depend partly on precision, and precision is often strongly influenced by sample size. As samples grow, smaller effects can become statistically detectable.

This is not a defect in statistics. It becomes a problem when researchers treat detectability as importance.

Cochrane's guidance on interpreting evidence makes this point directly: a small p-value can accompany an effect too trivial to produce an important benefit. Cochrane therefore advises review authors to focus interpretation on effect estimates and confidence intervals rather than relying on a binary distinction between statistically significant and non-significant findings.

The reverse problem occurs with small samples. A potentially consequential effect may fail to cross a significance threshold simply because the estimate is imprecise.

Effect estimates answer the question significance tests cannot

If your substantive question is “How much does X change Y?”, the estimated magnitude of the effect needs to be central to the analysis.

Depending on the research question and outcome, researchers might report mean differences, standardized mean differences, risk differences, risk ratios, odds ratios, correlations, regression coefficients, rate ratios, or other appropriate effect measures.

The specific metric matters less than the underlying principle: report the quantity that represents the effect you want to understand and interpret its magnitude in context.

An effect estimate without context can still be difficult to interpret. A standardized effect of 0.30, for example, does not become “small” in every field merely because a conventional label says so. Whether an effect matters depends on the outcome, baseline conditions, costs, risks, feasibility, available alternatives, and what differences are consequential in that domain.

Confidence intervals help reveal what remains plausible

A point estimate is not known with perfect precision. Confidence intervals provide information about sampling uncertainty around that estimate under the statistical procedure used.

Cochrane recommends interpreting point estimates together with confidence intervals because the interval can reveal whether the data remain compatible with importantly different conclusions.

Suppose a study estimates an improvement of four points, with a confidence interval extending from one to seven points. If a three-point improvement would be substantively important, the interval includes both effects below and above that threshold. The study therefore leaves uncertainty about whether the true effect is large enough to matter.

This is more informative than simply observing whether the interval excludes zero.

What the literature reports What it establishes What may remain unanswered
p <.05 The observed data meet a chosen significance criterion under the specified analysis How large and consequential the effect is
p >.05 The chosen significance criterion was not met Whether important effects remain plausible because of imprecision
Effect estimate without uncertainty An estimated magnitude How precisely that magnitude is known
Effect estimate and confidence interval Magnitude plus sampling uncertainty Whether the plausible effects are meaningful in the relevant context
Very precise tiny effect A small effect estimated with considerable precision Whether the difference matters enough to influence understanding or action

“Non-significant” does not mean “no effect”

A study that fails to cross a conventional significance threshold has not necessarily demonstrated the absence of an effect.

Cochrane specifically warns against confusing lack of evidence of an effect with evidence of a lack of effect. A wide confidence interval might include no effect, a small effect, and a substantial effect simultaneously.

The appropriate interpretation depends on the estimate and its uncertainty.

If researchers genuinely want to establish that effects are sufficiently small to be considered unimportant, approaches designed around equivalence, non-inferiority, or predefined ranges of practical equivalence may sometimes be more suitable than conventional null-hypothesis significance testing. Which approach is appropriate depends on the substantive question.

A meaningful difference must be justified, not invented after seeing the results

Once researchers recognize that statistical significance is insufficient, another temptation appears: declaring whatever observed effect they obtained “practically significant.”

That solves nothing.

A threshold for meaningful change should ideally be justified independently of the observed result. Depending on the field, evidence may come from established minimal important differences, stakeholder judgments, validated benchmarks, policy thresholds, cost-benefit considerations, prior empirical work, or theoretically defensible criteria.

In some areas, no accepted threshold exists. Researchers should say so rather than manufacture one.

Watch Out

Replacing an arbitrary p <.05 threshold with an equally arbitrary “meaningful effect” threshold does not solve the interpretive problem. The substantive criterion itself needs justification.

Meaningfulness may differ across stakeholders and contexts

A modest effect can be important when an intervention is inexpensive, safe, scalable, and applied to millions of people. The same effect may be unattractive when achieving it requires substantial cost, risk, workload, or displacement of a better alternative.

Likewise, what counts as meaningful to an individual participant may differ from what matters to an institution, policymaker, clinician, teacher, or researcher testing a theory.

This is why practical significance should not be reduced to a universal table of effect-size labels.

The outcome itself must also matter

A large effect on a trivial or indirect outcome does not automatically become consequential.

If studies detect substantial changes in engagement but the real decision concerns durable learning, the unresolved issue may be that researchers have measured an outcome that does not adequately answer the substantive question.

Magnitude and outcome relevance therefore need to be considered together.

Statistical significance can hide the real research gap

Imagine 40 studies demonstrating statistically significant effects of an intervention. A conventional review might conclude that another study is unnecessary.

But suppose most papers report little about magnitude, confidence intervals are rarely interpreted substantively, and nobody has established whether the observed differences exceed a threshold that matters in practice.

The unanswered question is no longer simply whether the intervention has an effect. It is whether the effect is consequential enough to justify the conclusions or decisions being made.

This is a good example of why research uncertainty can persist even when studies themselves are not missing.

04 · A Practical Example

When a Highly Significant Improvement May Still Be Too Small to Matter

Hypothetical Example

Does an AI tutoring system improve examination performance?

Imagine a hypothetical randomized study involving 10,000 university students. Students using an AI tutor score an average of 80.3 on the final assessment, compared with 80.0 among students in the comparison condition.

Statistical result Because the sample is very large and outcome variability is sufficiently low, the 0.3-point difference produces a very small p-value.
Weak conclusion The AI tutor significantly improves academic performance and is therefore educationally effective.
Missing question Is a 0.3-point improvement large enough to matter educationally, particularly given the cost, workload, alternatives, and other consequences of implementing the system?
Better interpretation Report the estimated difference and its uncertainty, then evaluate whether the plausible effect is large enough to be consequential using a defensible substantive criterion.

Now reverse the situation. Suppose a study of 80 students estimates a four-point improvement but with a wide confidence interval that includes both trivial and substantial effects. A p-value above.05 would not establish that the intervention has no worthwhile effect. It would indicate that the study has not estimated the effect precisely enough to distinguish among those possibilities.

05 · What Researchers Often Get Wrong

Common Mistakes When Moving Beyond Statistical Significance

Misconception

Statistically Significant Means Important

No. Statistical significance does not measure effect magnitude or substantive importance. A tiny effect can produce a very small p-value when estimated precisely enough.

Misconception

Non-Significant Means There Is No Effect

No. A result that does not cross a significance threshold may still be compatible with effects ranging from negligible to consequential. The effect estimate and its uncertainty are necessary for interpretation.

Misconception

A Larger Effect Size Is Always More Important

Not necessarily. Importance depends on the outcome, scale, baseline conditions, costs, harms, feasibility, alternatives, and context. Even a modest effect may be consequential when applied broadly or achieved cheaply, while a larger effect may not justify substantial burdens.

Misconception

Cohen's Conventional Labels Tell You Whether an Effect Matters

Generic effect-size conventions can sometimes provide descriptive reference points, but they should not automatically determine substantive importance. What counts as consequential should be justified within the research domain and decision context.

Misconception

Confidence Intervals Matter Only Because They Show Statistical Significance

Their more useful role is showing the uncertainty around an estimate. Researchers can examine whether the interval includes effects that would lead to substantively different conclusions, rather than using it merely to check whether zero is excluded.

06 · What This Means for You

Build the Research Question Around the Magnitude That Would Matter

When a literature is saturated with significance tests, reconstruct the question in quantitative and substantive terms.

Do not ask only whether X affects Y. Ask how much X changes Y, how precisely that difference is known, and what magnitude would alter the relevant scientific or practical interpretation.

A simple decision framework

If previous studies report mostly p-values
Recover or examine effect estimates and uncertainty wherever possible rather than counting significant findings.
If effects are precisely estimated but very small
Ask whether their magnitude is consequential for the relevant theory, practice, population, or decision.
If potentially important effects remain imprecise
Design research capable of narrowing uncertainty around the magnitude that matters.
If no meaningful-effect threshold exists
Develop a transparent substantive justification rather than choosing a threshold after seeing the data.
If the outcome itself has questionable relevance
Resolve the outcome problem before debating whether the observed effect is large enough.

The research contribution then becomes clearer. Instead of trying to obtain one more significant result, you are attempting to determine whether the plausible effect is large enough to matter.

07 · A Quick Checklist

Before Treating Statistical Significance as an Answer, Check the Effect

When reviewing statistically focused research, check:
What effect magnitude do the studies actually estimate?
How precise are those estimates, and what substantively different effects remain compatible with the uncertainty?
Are authors interpreting p-values as though they measure effect size or importance?
Could large samples be making trivial effects statistically detectable?
Could small samples be leaving consequential effects too imprecise to distinguish?
What magnitude of effect would actually matter for this outcome and decision?
Is that meaningful-effect criterion empirically, theoretically, or practically justified?
Would a new study narrow uncertainty about meaningful effects rather than merely seek another p-value below.05?
08 · Frequently Asked Questions

Questions About Statistical and Meaningful Effects

Does p <.05 mean an effect is important?

No. A p-value does not measure the size or importance of an effect. Interpret the estimated magnitude, its uncertainty, the outcome being measured, and the substantive context.

Does p >.05 mean there is no effect?

No. It means the result did not meet the specified significance threshold. The estimate and confidence interval may still be compatible with meaningful effects, particularly when evidence is imprecise.

Is effect size the same as practical significance?

No. An effect size quantifies magnitude on a particular scale. Practical or substantive importance concerns whether that magnitude matters in context. Interpreting importance may require considering benefits, harms, costs, feasibility, baseline conditions, alternatives, and stakeholder priorities.

What is a meaningful effect size?

There is no universal value. Meaningfulness depends on the construct, outcome, field, decision, population, and consequences. Where validated minimal important differences or other domain-specific thresholds exist, they may help; otherwise the criterion should be justified transparently.

Should researchers stop reporting p-values?

The more important issue is avoiding interpretations that p-values cannot support. The American Statistical Association cautions against using statistical significance as a substitute for scientific reasoning, while Cochrane recommends emphasizing effect estimates and confidence intervals rather than binary significance labels.

Can a very small effect still matter?

Yes. A small individual effect may be consequential when the outcome is important, the intervention is inexpensive or scalable, exposure is widespread, or effects accumulate. Context determines whether magnitude is meaningful.

09 · The Bottom Line

The Important Question Is Often How Much, Not Merely Whether

The Bottom Line

A literature focused on statistical significance can leave the central question unanswered because crossing a p-value threshold does not establish how large, precise, or meaningful an effect is.

Examine effect estimates and their uncertainty, then interpret them against a defensible substantive criterion. A useful next study should help determine whether the effect is large enough to matter, not merely whether another analysis can produce the label “statistically significant.”

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes