01 · The Question
Does P > 0.05 Mean the Study Found Nothing?
A paper reports P = 0.18. The authors call the result nonsignificant. Perhaps they go further and say there was "no effect," "no difference," or "no association."
Those statements are not equivalent. A conventional hypothesis test that fails to reject a null hypothesis has not automatically demonstrated that the null hypothesis is true. Sometimes the study is simply too imprecise to distinguish an important effect from no effect. In other cases, however, the data may genuinely provide useful evidence that effects large enough to matter are unlikely.
The key is to stop treating "nonsignificant" as the scientific conclusion and examine what the effect estimate and its uncertainty actually allow you to say.
03 · What You Need to Know
Not Significant Does Not Mean No Effect
What a nonsignificant hypothesis test actually tells you
In conventional null-hypothesis significance testing, researchers specify a null hypothesis and calculate a test statistic and P-value. If the P-value does not cross the prespecified significance threshold, such as 0.05, they fail to reject the null hypothesis.
That is not the same as proving the null hypothesis.
A result may fail to reach statistical significance because the true effect is small, because the data are noisy, because the sample provides insufficient information, or because the estimated effect happens to be closer to the null in this particular sample. The P-value alone does not distinguish these possibilities.
Failure to detect an effect
The analysis did not provide sufficient evidence to reject the specified null hypothesis at the chosen threshold.
Evidence that meaningful effects are absent
The data are sufficiently precise to rule out effects that were defined as substantively important under an appropriate inferential procedure.
Confidence intervals reveal whether the result is inconclusive or genuinely informative
Cochrane recommends focusing interpretation on effect estimates and their confidence intervals rather than dividing findings mechanically into statistically significant and nonsignificant categories. This is particularly important for results that do not cross a significance threshold.
Consider two hypothetical studies that both report an estimated mean difference of 1 point and P > 0.05.
Result
95% confidence interval
What the interval suggests
Study A
-12 to 14 points
Large negative and positive effects remain compatible with the data
Study B
-0.5 to 2.5 points
Effects far from zero are much less compatible with the data
Both results might receive the label "not statistically significant," but they contain very different amounts of information. Study A leaves enormous uncertainty. Study B estimates the effect much more precisely.
This is why P-values, effect sizes, and confidence intervals should be read together .
Ask which effects the study can reasonably rule out
Suppose researchers agree that differences smaller than 3 points would not be important for the decision at hand. If a study estimates a difference of 0.5 points with a 95% confidence interval from -0.8 to 1.8, the entire interval lies within that range of effects considered unimportant.
That is potentially much more informative than merely saying P > 0.05. The study has estimated the effect with enough precision that effects exceeding the specified threshold are not supported by the interval under the analysis.
By contrast, if the confidence interval runs from -8 to 9, important effects in either direction remain plausible. Calling that result "no difference" would conceal the study's uncertainty.
Equivalence testing asks a more appropriate question when absence matters
A conventional null-hypothesis significance test is usually designed to test an exact null such as a difference of zero. Failing to reject that null is not designed to demonstrate that effects large enough to matter are absent.
Equivalence testing reverses the problem. Researchers specify bounds representing effects that would be considered meaningfully different from zero. A procedure such as the two one-sided tests approach can then test whether effects at least as extreme as those bounds can be rejected.
Lakens and colleagues emphasize that a study can be nonsignificant in a conventional test yet also fail an equivalence test. In that situation, the result is not evidence for a meaningful effect and not adequate evidence for its absence. It is simply inconclusive.
The meaningful bounds should have substantive justification and, ideally, be determined without tailoring them after seeing the observed result.
A nonsignificant result can be scientifically useful even when it is inconclusive
"Informative" does not always mean that the study settles the question.
An imprecise nonsignificant result may still contribute data to a later meta-analysis, reveal that an anticipated effect is probably smaller than expected, expose problems with measurement or feasibility, or help refine the design of subsequent studies. The evidential contribution should simply not be exaggerated.
A single underpowered study may provide weak evidence on its own while still adding information when combined appropriately with other comparable studies.
Statistical power matters, but post hoc labels can mislead
Low information or poor precision can make meaningful effects difficult to distinguish from the null. This is one reason recognizing an underpowered study from the published paper can matter when interpreting a nonsignificant finding.
However, simply calculating "observed power" from the observed effect and then declaring the result uninformative usually adds little. The observed P-value and conventional observed power are mathematically closely related. A more useful appraisal focuses on the study's design, anticipated effect sizes, sample size or event count, and especially the width of the confidence interval.
Nonsignificant results can contradict strong claims
Suppose a theory predicts a very large effect. A well-designed study estimates an effect near zero with a narrow confidence interval that excludes effects remotely close to the theoretical prediction. Even if the conventional test against an exact zero happens not to be significant, the result may be highly relevant to evaluating that strong prediction.
This illustrates why the scientific hypothesis and the statistical null hypothesis should not be treated as interchangeable. A theory rarely predicts merely "anything other than exactly zero."
Do not compare significance labels instead of effects
Another common error occurs when one subgroup shows P < 0.05 and another shows P > 0.05. Researchers may conclude that the effect exists in one subgroup but not the other.
That inference does not follow from the two significance labels. A direct test of the difference or interaction between groups is generally required to evaluate whether the effects differ. This deserves particular scrutiny when a subgroup finding was not prespecified .
Absence claims require especially careful wording
"We found no statistically significant difference" describes a test outcome. "There is no difference" is a claim about the underlying phenomenon. The second statement generally requires more evidence than the first.
When the data remain compatible with important effects, use language that preserves the uncertainty. When sufficiently precise evidence or an appropriate equivalence analysis rules out effects large enough to matter, a stronger conclusion about the absence of a meaningful effect may become defensible.
Watch Out
Do not interpret P > 0.05 as evidence that two treatments, groups, exposures, or conditions are equivalent. Failure to detect a difference and evidence of equivalence are distinct statistical conclusions.
04 · A Practical Example
Two Nonsignificant Results Can Tell Very Different Stories
Hypothetical Example
Two studies evaluate the same educational intervention
Suppose a 5-point improvement on a 100-point assessment would be considered educationally meaningful. Two independent studies both fail to reject an exact zero-effect null at the conventional 0.05 level.
Study A: inspect the estimate. The estimated improvement is 2 points, with a 95% confidence interval from -7 to 11.
Interpret Study A. The interval includes no effect, meaningful benefit, and potentially meaningful harm. The nonsignificant result has not resolved the practical question.
Study B: inspect the estimate. The estimated improvement is 0.4 points, with a 95% confidence interval from -1.2 to 2.0.
Interpret Study B. The interval is much narrower and does not extend to the prespecified 5-point benefit threshold in either direction. Under the assumptions of the analysis, the data provide considerably more information against effects of that magnitude.
Compare the studies. Calling both findings simply "nonsignificant" erases their most important difference. One study is highly uncertain; the other provides a comparatively precise estimate close to zero.
If the researchers' explicit goal were to establish equivalence within specified bounds, an appropriate equivalence analysis would provide a more direct inferential test than relying only on the conventional P-value.
06 · What This Means for You
Ask What the Study Has Learned, Not Which Side of 0.05 It Landed On
When a paper reports a nonsignificant finding, move directly to the effect estimate and confidence interval. Your goal is to determine which scientifically meaningful possibilities remain compatible with the data.
A simple decision framework
If the confidence interval is wide and includes important effects in either direction
Treat the result as inconclusive rather than as evidence of no effect.
If the estimate is near the null and the interval is narrow enough to exclude effects large enough to matter
The result may provide useful evidence against substantively important effects, subject to the validity of the study and analysis.
If the research question specifically concerns equivalence or absence of a meaningful effect
Look for an equivalence test or another inferential method designed to evaluate that question rather than relying only on nonsignificance.
If the study is small and the interval is highly imprecise
Avoid strong absence claims and consider whether additional evidence or synthesis across studies is needed.
A useful nonsignificant result narrows uncertainty. The critical question is how much uncertainty remains after seeing the data and whether the remaining plausible effects would lead you to materially different scientific or practical conclusions.
07 · A Quick Checklist
Before Interpreting a Nonsignificant Result
Check whether:
You have examined the effect estimate rather than stopping at P > 0.05.
A confidence interval or another appropriate measure of uncertainty is reported.
The interval is narrow enough to distinguish important effects from effects too small to matter.
Any threshold for a meaningful effect has a substantive justification.
The study had enough information, observations, or events to estimate the effect with useful precision.
The authors have avoided equating failure to reject the null with proof of no effect.
Claims of equivalence are supported by an appropriate analysis rather than nonsignificance alone.
The study design and analysis are credible enough for either presence or absence claims to be meaningful.
09 · The Bottom Line
Nonsignificant Does Not Mean Uninformative, and It Does Not Mean No Effect
The Bottom Line
A nonsignificant result can be informative when the estimate is precise enough to narrow the range of plausible effects, but P > 0.05 alone does not demonstrate that an effect is absent. Interpret the effect estimate and its uncertainty rather than treating the significance label as the conclusion.
If important effects remain inside a wide confidence interval, the appropriate interpretation may simply be uncertainty. If meaningful effects can be excluded with adequate precision, particularly using a method designed for that purpose, the result can provide useful evidence about how small the effect is likely to be.
11 · Cite this Guide
How to Cite This Guide
This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.
Recommended (Field Guide)
APA
MLA
Chicago
Copy Citation