01 · The Question
Does failing to detect an effect mean there is no effect?
A study compares two groups and obtains p =.12. The Results say that the difference was not statistically significant. The Discussion says there was “no evidence of a difference.” Then the Conclusion says the intervention “has no effect.”
Those statements are not necessarily equivalent.
A study can fail to provide convincing evidence for an effect because the true effect is small or absent. But the same outcome can occur when the data are too imprecise to distinguish no effect from an effect large enough to matter. Critical reading requires you to determine which situation the evidence actually supports.
03 · What You Need to Know
Why “not statistically significant” does not mean “no effect”
Start with three different statements
Statement
What it means
What it does not automatically mean
No statistically significant difference was detected
The analysis did not cross the specified statistical-significance threshold.
The true effect is zero.
The evidence is inconclusive
The data remain compatible with materially different possibilities.
The alternatives have been shown to be equivalent.
Evidence supports absence of a meaningful effect
The analysis provides information capable of ruling out effects considered important under the specified criterion and assumptions.
Every conceivable nonzero effect has been disproved.
The second and third situations are particularly easy to confuse. A study can be inconclusive precisely because it cannot tell you whether an important effect exists.
Read the effect estimate before the significance label
Suppose a trial estimates that an intervention improves a score by 3 points, with a 95% confidence interval from -2 to 8 points.
If the conventional null value for the difference is zero, this interval includes zero. A conventional superiority test may therefore be statistically non-significant.
But the interval also includes a possible improvement of 8 points. If an 8-point improvement would matter substantially, the data have not ruled out an important benefit.
Calling the result “no effect” would erase that uncertainty.
CONSORT 2025 specifically warns against interpreting a non-significant result as equivalence and emphasizes confidence intervals because they can show that a statistically non-significant result remains compatible with an important effect.
A p-value does not estimate how large the effect is
The American Statistical Association cautions that a p-value or statistical-significance designation does not measure effect size or the importance of a result. Scientific conclusions should not be based only on whether a p-value passes a threshold.
This matters because p-values depend partly on the amount and variability of information in the study. Similar estimated effects can produce different p-values in studies with different precision.
Therefore:
p >.05 does not translate mathematically into “effect = 0.”
Nor does p <.05 mean the effect is large or important. This is one reason to keep statistical significance separate from practical importance .
Wide uncertainty makes “no effect” especially problematic
Imagine two studies with the same point estimate of zero.
Study
Estimated difference
95% confidence interval
What the interval suggests
Study A
0
-1 to 1
The data are relatively precise around a small range.
Study B
0
-12 to 12
The data remain compatible with substantial effects in either direction.
Both point estimates are zero. Both studies might produce a non-significant test against a null difference. Yet they provide very different information.
Study B tells you far less about whether an important effect is absent. This is why reading only “not significant” discards information you need.
Look for the linguistic shift across the paper
The error often appears gradually:
Results: “The difference was not statistically significant.”
Discussion: “We found no evidence that the intervention improved performance.”
Conclusion: “The intervention does not improve performance.”
The first statement reports the result of a statistical decision rule. The second may be a reasonable summary if appropriately understood. The third makes a stronger claim about absence.
Place these sentences next to each other. If certainty increases while the underlying evidence remains unchanged, ask what justifies the stronger conclusion.
“No evidence of X” can itself be misunderstood
Even the phrase “no evidence of an effect” needs care. It can sound as though the study produced affirmative evidence that no effect exists.
Sometimes a more informative statement is simply to report the estimate and uncertainty:
“The estimated difference was 1.4 points, with a 95% confidence interval from -2.1 to 4.9 points.”
This lets readers see what the data remain compatible with instead of compressing the entire result into a binary label.
Evidence of absence requires a meaningful definition of “absence”
In most applied research, the scientifically interesting question is rarely whether the true effect is exactly zero to infinite precision.
A more useful question may be whether any remaining effect is small enough to be considered practically unimportant.
That requires a margin or criterion. How small would a difference need to be before you would treat the interventions as practically equivalent? What difference would matter educationally, clinically, economically, or otherwise?
Without such a criterion, “no meaningful effect” can become an undefined claim.
Equivalence testing asks a different question from ordinary superiority testing
A conventional superiority analysis typically asks whether the data provide sufficient evidence of a difference under a specified null hypothesis. Failing to demonstrate superiority does not reverse the test and demonstrate equivalence.
An equivalence design instead asks whether the effect is sufficiently bounded within a prespecified equivalence region. The equivalence margin should have a substantive justification rather than being selected after the results are known merely to obtain the desired conclusion.
Similarly, a non-inferiority trial asks whether an intervention can be shown not to be worse than a comparator by more than a prespecified margin. CONSORT 2025 treats superiority, equivalence, and non-inferiority as distinct trial frameworks that should be explicitly identified.
Therefore, if a conventional superiority trial obtains p >.05 and the authors declare the treatments “equivalent,” check whether an actual equivalence analysis was planned and conducted.
Sample size matters through precision, but “underpowered” is not a complete interpretation
A small sample can produce imprecise estimates, making it difficult to distinguish meaningful effects from no effect. But simply labeling a study “underpowered” after seeing a non-significant result is not enough.
Inspect the effect estimate and its uncertainty directly. The confidence interval can show which effect sizes remain compatible with the data.
A study with a smaller-than-planned sample might still provide useful information if its estimate is sufficiently precise for the question at hand. Conversely, a nominally large study can leave important uncertainty for rare outcomes or noisy measurements.
Precision is the evidential issue you ultimately care about.
Evidence of absence is possible
The slogan “absence of evidence is not evidence of absence” is a warning, not a rule that absence can never be supported.
Studies can provide evidence against effects of a specified meaningful magnitude. Equivalence tests, well-designed non-inferiority analyses, sufficiently precise confidence intervals, and some Bayesian approaches can all contribute to such conclusions when used appropriately.
The critical distinction is between failing to detect something and obtaining evidence that meaningfully constrains how large it could be.
Watch Out
Do not respond to every non-significant finding by insisting that “there could still be an effect” without considering magnitude. With sufficiently precise evidence, remaining plausible effects may be too small to matter for the research question. Critical appraisal should preserve uncertainty, not manufacture it.
06 · What This Means for You
Replace binary significance reading with an uncertainty check
Whenever a paper claims that something “does not work,” “has no effect,” “makes no difference,” or is “equivalent,” do not begin with the p-value. Begin with the inferential question and the range of effects still compatible with the evidence.
A simple null-result framework
If the estimate is imprecise and includes important effects
Treat the result as inconclusive rather than evidence that no meaningful effect exists.
If the estimate is precise and excludes effects large enough to matter
The study may provide useful evidence against a meaningful effect, provided the threshold and assumptions are defensible.
If the authors claim equivalence
Check whether equivalence margins were justified and an appropriate equivalence design and analysis were used.
If the authors claim non-inferiority
Verify the prespecified non-inferiority margin, analysis, confidence interval, and other design requirements relevant to that claim.
If the conclusion simply says “no effect” after p >.05
Return to the estimate and uncertainty before accepting the stronger wording.
This is another reason to keep the numerical result separate from the authors’ narrative . A binary sentence can conceal a great deal of uncertainty.
It also prevents an inconclusive finding from becoming one of the claims introduced more strongly in the Conclusion than the Results support .
07 · A Quick Checklist
Did “not detected” quietly become “does not exist”?
When authors report no effect or no difference, check:
Find the effect estimate rather than relying on the significance label.
Examine the confidence interval or other representation of uncertainty.
Determine whether effects large enough to matter remain compatible with the data.
Identify the substantive threshold that defines a negligible or meaningful effect.
Check whether the study was designed for superiority, equivalence, non-inferiority, or another inferential goal.
Do not treat p >.05 as proof that the null hypothesis is true.
Compare the wording in the Results with the stronger wording, if any, in the Discussion and Conclusion.
Distinguish an inconclusive study from a study that meaningfully rules out important effects.
11 · Cite this Guide
How to Cite This Guide
This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.
Recommended (Field Guide)
APA
MLA
Chicago
Copy Citation