Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Should You Focus on P-Values, Effect Sizes, Confidence Intervals, or All Three?

P-values, effect sizes, and confidence intervals answer different questions. Good interpretation usually emphasizes the estimated effect and its uncertainty while using the P-value as supplementary evidence rather than the entire conclusion.

385
P-Values, Effect Sizes, or Confidence Intervals? Guide 385 of 899
01 · The Question

Which Statistical Result Deserves Your Attention?

A paper reports a treatment effect of 4.2 points, a 95% confidence interval of 0.3 to 8.1, and P = 0.034. Which number tells you whether the finding matters?

None of them answers that question completely by itself. The effect estimate tells you how large the observed difference is. The confidence interval tells you about uncertainty around that estimate. The P-value summarizes how incompatible the observed data are with a specified statistical model, typically one involving a null hypothesis.

The mistake is treating these quantities as competitors and choosing one number to carry the entire interpretation. They answer related but distinct questions.

02 · The Short Answer

Usually Examine All Three, but Do Not Give Them Equal Interpretive Weight

In Brief

When P-values, effect estimates, and confidence intervals are available, interpret them together, with particular attention to the magnitude of the estimated effect and the uncertainty around it. A P-value can contribute information about compatibility with a null hypothesis, but it does not tell you how large, important, or practically meaningful an effect is.

The exact balance depends on the research design and inferential framework. In many conventional analyses, the effect estimate and confidence interval provide more directly useful information for substantive interpretation than a binary declaration that P is above or below 0.05.

03 · What You Need to Know

Each Quantity Answers a Different Statistical Question

The effect estimate asks: how large is the observed effect?

An effect estimate describes the magnitude and direction of an observed relationship, difference, or association. Depending on the research question, this might be a mean difference, risk difference, risk ratio, odds ratio, hazard ratio, correlation coefficient, regression coefficient, standardized mean difference, or another estimand.

Suppose an intervention reduces a mean outcome by 0.2 points. Knowing that P < 0.001 does not tell you whether a reduction of 0.2 points matters. That requires understanding the effect estimate in its substantive context.

Effect magnitude may be expressed on an absolute scale or a relative scale. The choice can materially affect interpretation. A relative reduction of 50%, for example, could represent a decrease in risk from 20% to 10% or from 0.02% to 0.01%. The relative change is the same, but the absolute consequences are very different.

The confidence interval asks: how uncertain is the estimate?

A point estimate is only one estimate obtained from a sample. A confidence interval expresses statistical uncertainty around that estimate under the assumptions of the procedure used to construct it.

For a conventional 95% frequentist confidence interval, the formal interpretation concerns the long-run performance of the interval-generating procedure: if the procedure were repeated across many comparable samples, 95% of the resulting intervals would contain the true parameter under the model assumptions. It is therefore imprecise to say that there is a 95% probability that the fixed true parameter lies inside one particular frequentist interval.

In practical appraisal, interval width is especially informative. A narrow interval suggests greater precision, while a wide interval may leave several substantively different effects compatible with the data. Cochrane explicitly recommends considering point estimates together with their confidence intervals when interpreting intervention effects.

The P-value asks a narrower question

A P-value is calculated relative to a statistical model and hypothesis. In a conventional null-hypothesis significance test, it represents the probability, assuming the null hypothesis and other model assumptions used in the calculation, of obtaining a result at least as incompatible with that hypothesis as the observed result.

That definition rules out several common interpretations. A P-value of 0.03 does not mean there is a 3% probability that the null hypothesis is true. It does not mean there is a 97% probability that the research hypothesis is correct. It does not tell you the probability that the result occurred "by chance." And it does not measure the size or practical importance of the effect.

Quantity Main question it helps answer What it does not tell you by itself
Effect estimate How large and in what direction is the observed effect? How precise the estimate is
Confidence interval What range of parameter values is compatible with the data and model at the stated confidence level? Whether every value in that range is substantively important
P-value How incompatible are the observed data with the specified null model, according to the chosen test? The size, importance, or probability that the hypothesis is true

Why P < 0.05 should not become the conclusion

The familiar 0.05 threshold is a convention, not a boundary separating real effects from unreal ones. Results with P = 0.049 and P = 0.051 provide nearly the same numerical evidence, yet dichotomizing them can make their interpretations sound dramatically different.

Cochrane guidance consequently advises against undue reliance on labels such as statistically significant and nonsignificant and recommends focusing interpretation on estimates of effect and confidence intervals, with exact P-values reported when used.

The American Statistical Association has likewise emphasized that scientific conclusions should not be based only on whether a P-value crosses a particular threshold and that statistical significance does not measure the size or importance of an effect.

A small P-value can accompany a tiny effect

P-values depend partly on precision, which is strongly influenced by sample size. With sufficiently large samples, very small effects can generate very small P-values.

Imagine a digital intervention that improves an outcome by 0.05 points on a 100-point scale. In an enormous dataset, the estimate might be extremely precise and yield P < 0.001. The evidence may be strong that the population effect is not exactly zero, while the magnitude remains practically negligible.

This is why a statistically significant result can still be unimportant and why very large samples can make trivial effects appear statistically impressive.

A large P-value does not demonstrate that there is no effect

The reverse mistake is equally common. P > 0.05 does not establish that the null hypothesis is true or that the groups are equivalent. A study may simply be too imprecise to distinguish among several plausible effects.

Suppose an estimated mean difference is 5 points with a 95% confidence interval from -2 to 12. The interval includes zero, but it also includes potentially meaningful positive effects. Describing this only as "no significant difference" discards important information about what the study can and cannot exclude.

This becomes especially important when evaluating whether a nonsignificant result is nevertheless informative or whether the study may be too imprecise to answer its main question convincingly.

Confidence intervals connect magnitude and uncertainty

One reason confidence intervals are so useful is that they force you to consider more than whether the null value is included. Ask what the entire interval permits.

For a mean difference, the null value is typically 0. For ratio measures such as a risk ratio or odds ratio, it is typically 1. Whether the interval crosses that null value relates to conventional statistical significance when the confidence interval and hypothesis test use corresponding methods. But stopping there wastes much of the information the interval contains.

Consider an estimated risk ratio of 0.80 with a 95% confidence interval from 0.77 to 0.83. That tells a different evidential story from a risk ratio of 0.80 with an interval from 0.45 to 1.42, even though the point estimates are identical. The second study leaves much greater uncertainty about the underlying effect.

Effect size does not automatically mean practical importance

The phrase "effect size" can create another shortcut. A standardized effect of a certain numerical magnitude is not automatically small, medium, or large in every discipline or application. Context determines whether an effect matters.

A seemingly small effect may be important when an intervention is inexpensive, scalable, low risk, and applied to millions of people. A numerically larger effect may matter little if it occurs on an outcome with minimal practical relevance or comes with substantial costs and harms.

Whenever possible, interpret effects using domain knowledge, meaningful units, baseline risk, clinically or practically important thresholds, costs, harms, and consequences rather than relying mechanically on generic magnitude labels.

All three quantities inherit the weaknesses of the underlying analysis

A beautifully narrow confidence interval around a biased estimate is still a precisely estimated biased result. A tiny P-value from an inappropriate model does not rescue the model. A large effect estimate generated by selective reporting remains vulnerable to selective reporting.

Statistical summaries should therefore be interpreted only after asking whether the design, measurement, analysis, and reporting are credible. This is also why successfully recalculating the reported statistics cannot by itself establish that the study's inference is valid.

Watch Out

Do not read a confidence interval merely as a more elaborate significance test. Asking only whether it crosses the null value turns a range of information about magnitude and precision back into the same binary P < 0.05 decision you were trying to move beyond.

04 · A Practical Example

How the Same Result Changes When You Read All Three Numbers

Hypothetical Example

An educational intervention improves examination scores

Suppose a randomized study compares a new learning intervention with usual instruction. The adjusted mean difference in the final examination score is 2.0 points on a 100-point scale, with a 95% confidence interval from 0.4 to 3.6 points and P = 0.015.

Read the P-value. P = 0.015 indicates that the observed result is relatively incompatible with the specified null model under the assumptions of the test. It does not tell you whether a 2-point improvement matters educationally.
Read the effect estimate. The estimated difference is 2 points. Now the substantive question becomes unavoidable: is a 2-point improvement meaningful on this assessment?
Read the confidence interval. The interval from 0.4 to 3.6 points shows the uncertainty around the estimate. Effects near the lower end may be negligible, while effects near the upper end might be more consequential, depending on the educational context.
Add substantive knowledge. Suppose researchers had good reason, established independently of the observed result, to regard a 5-point difference as the smallest effect that would justify the cost of implementing the intervention. The entire reported interval lies below that benchmark.
Interpret rather than label. The study provides evidence against an exact zero-effect null under its statistical model, but the estimated improvement and its confidence interval would not support the claim that the intervention achieves the stipulated 5-point threshold.

The example shows why "statistically significant" is not a sufficient interpretation. The P-value answers one inferential question, while the effect estimate, confidence interval, and substantive threshold address questions that are usually closer to the decision a researcher actually cares about.

05 · What Researchers Often Get Wrong

Common Mistakes When Reading Statistical Results

Misconception

P < 0.05 Means the Finding Is Important

No. Statistical significance does not measure substantive importance. A tiny effect can have a very small P-value when estimated with sufficient precision.

Misconception

P > 0.05 Means There Is No Effect

A nonsignificant test does not establish absence of an effect. Examine the point estimate and confidence interval to determine which effects remain compatible with the data and how much uncertainty remains.

Misconception

The P-Value Is the Probability That the Null Hypothesis Is True

It is not. A conventional P-value is calculated conditional on the null hypothesis and other model assumptions. It does not provide the posterior probability that the null hypothesis itself is true.

Misconception

If the 95% Confidence Interval Includes Zero, the Effect Must Be Zero

No. Inclusion of the null value means that the corresponding conventional test may not reject the null at the 0.05 level when the procedures correspond. The interval may simultaneously include effects that are substantively important in either direction.

Misconception

A Standardized Effect Size Has the Same Meaning Everywhere

Generic labels such as small, medium, and large can be convenient summaries, but substantive importance depends on the outcome, population, intervention, baseline conditions, costs, risks, and disciplinary context.

Misconception

A Narrow Confidence Interval Means the Study Is High Quality

A narrow interval indicates statistical precision under the analysis used. It does not rule out bias, confounding, measurement problems, selective reporting, inappropriate modeling, or other threats to validity.

06 · What This Means for You

Read the Estimate First, Then Its Uncertainty, Then the P-Value in Context

When critically appraising a conventional statistical result, a useful reading sequence is to begin with the estimated effect and ask what it means in the real units or substantive context of the study. Then examine the confidence interval to understand how precisely that effect has been estimated. Finally, use the P-value, when reported, as additional information about the specified statistical test rather than as the verdict on the study.

A simple interpretation framework

If the effect estimate is substantively important and the interval is reasonably precise
Ask whether the design and analysis are credible enough for that estimate to support the claimed inference.
If P is small but the estimated effect is trivial
Do not confuse strong evidence against an exact null with evidence of practical importance.
If P is large and the confidence interval is wide
Treat the result as imprecise rather than concluding automatically that there is no effect.
If the confidence interval excludes effects large enough to matter
That may be substantively informative even if the usual significance label does not capture the point you care about.

The hierarchy is not absolute. Some research questions use different inferential frameworks, and not every study reports all three quantities. The general principle is to extract as much information as the analysis legitimately provides without allowing one threshold to replace scientific interpretation.

07 · A Quick Checklist

How to Read a Statistical Result Without Stopping at P < 0.05

When interpreting a statistical result, check:
What effect or parameter is actually being estimated?
What is the magnitude and direction of the effect in substantively meaningful terms?
Is an absolute effect available as well as a relative or standardized effect where that would improve interpretation?
How wide is the confidence interval, and what substantively different effects does it include?
What null hypothesis and statistical model generated the P-value?
Are you interpreting the P-value as evidence about compatibility with the null model rather than as the probability that a hypothesis is true?
Have you separated statistical significance from practical, clinical, educational, or theoretical importance?
Are the design and analysis credible enough for any of these statistical summaries to support the claimed inference?
08 · Frequently Asked Questions

Questions About P-Values, Effect Sizes, and Confidence Intervals

Should I ignore P-values completely?

No. P-values can provide information about the compatibility of data with a specified statistical model or null hypothesis. The problem is using them as a standalone measure of effect magnitude, practical importance, study quality, or truth.

Is an effect size always better than a P-value?

They answer different questions. An effect estimate communicates magnitude, while a P-value relates the data to a specified hypothesis and model. For substantive interpretation, magnitude is usually indispensable, but its uncertainty and the credibility of the underlying analysis also matter.

Is a confidence interval better than both?

A confidence interval is especially informative because it displays uncertainty around an estimate, but it does not independently tell you whether the design is unbiased, whether the effect is important, or whether model assumptions are appropriate. It should be interpreted with the point estimate and substantive context.

Does a 95% confidence interval mean there is a 95% probability that the true value is inside it?

Not under the conventional frequentist interpretation. The 95% refers to the long-run coverage of the interval-generating procedure under its assumptions. Probability statements about the parameter itself require a different inferential framework, such as an appropriately specified Bayesian analysis.

What if the confidence interval barely crosses zero?

Do not turn that fact into a binary conclusion. Examine the estimate, the entire interval, and the substantive importance of the values it contains. An interval from -0.1 to 5.0 tells a very different story from one spanning -20 to 25 even though both include zero.

What is more important: statistical significance or practical significance?

They concern different questions. Statistical significance describes a result relative to a statistical testing procedure, while practical significance concerns whether the magnitude matters in the relevant real-world or theoretical context. A useful interpretation should not substitute the former for the latter.

Can a large effect size still be uncertain?

Yes. A study can produce a large point estimate with a very wide confidence interval, especially when the sample or number of events is small. The estimate may therefore be compatible with effects ranging from modest to extremely large, or sometimes even with effects in the opposite direction.

09 · The Bottom Line

Do Not Ask One Statistic to Answer Three Different Questions

The Bottom Line

Use effect estimates to understand magnitude, confidence intervals to understand statistical uncertainty, and P-values, when relevant, as evidence about compatibility with a specified hypothesis and model. In most conventional research appraisal, the estimate and its uncertainty deserve more attention than whether the P-value happens to cross 0.05.

None of these quantities establishes importance or validity on its own. Their interpretation still depends on study design, analytical assumptions, bias, substantive context, and the size of an effect that would actually matter for the research question.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes