Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Is Optional Stopping, and Why Can It Inflate False Positives?

Optional stopping occurs when accumulating results influence whether researchers continue collecting data. Under ordinary repeated significance testing, stopping when p < .05 can raise the false-positive rate above the nominal 5%.

495
Optional Stopping and False Positives Guide 495 of 530
01 · The Question

Why Does Checking the Data More Than Once Change the False-Positive Risk?

Suppose you begin with 40 participants and run your planned analysis. The result is not statistically significant, so you collect another 20 and test again. Still nothing. You add another 20, repeat the analysis, and continue until either p falls below.05 or you reach your maximum sample size.

You have not fabricated observations or altered the statistical test. Each individual analysis may look perfectly ordinary.

Yet the overall procedure is no longer equivalent to performing one fixed-sample test at the end. By repeatedly giving random variation opportunities to cross the significance threshold and stopping when it does, you can increase the probability of declaring an effect when no genuine effect exists.

02 · The Short Answer

Optional Stopping Makes the Stopping Rule Depend on the Data

In Brief

Optional stopping occurs when researchers use accumulating data or interim results to decide whether to stop or continue collecting observations. Under ordinary repeated unadjusted frequentist significance testing, repeatedly testing and stopping when p <.05 can increase the probability of a false-positive conclusion above the nominal significance level.

This does not mean all data-dependent stopping is statistically invalid. Properly designed sequential and adaptive methods explicitly incorporate interim analyses and stopping rules into the inferential procedure. The problem is treating an unplanned sequence of repeated tests as though only one fixed-sample test had occurred.

03 · What You Need to Know

How Optional Stopping Changes the Statistical Experiment

The Stopping Rule Is Part of the Procedure

Consider two studies that both finish with 100 participants.

Study A planned from the beginning to collect 100 observations and perform one hypothesis test afterward.

Study B began with 40 observations, tested the hypothesis, and then repeatedly added participants whenever the result remained nonsignificant. It happened to reach significance when the sample reached 100, so collection stopped.

The final sample sizes are identical. The procedures that produced them are not.

Study B gave the null hypothesis several opportunities to be rejected. The decision to reach 100 observations depended partly on what happened at earlier sample sizes.

Why Repeated Opportunities Can Inflate False Positives

Under a conventional fixed-sample frequentist test conducted at an alpha level of.05, the testing procedure is designed so that, under the relevant assumptions and when the null hypothesis is true, the long-run probability of rejection is 5%.

Now imagine repeatedly applying an ordinary.05 test to accumulating data and stopping the first time significance appears.

At one interim look, random sampling variation may push the result below.05. If it does, the study stops and records a “positive” result. If it does not, the researcher gets another opportunity later.

Those repeated opportunities are correlated because later datasets contain earlier observations, so you cannot calculate the overall false-positive probability by simply multiplying the number of looks by 5%. Nevertheless, under ordinary unadjusted repeated testing, the overall probability can exceed 5%.

A Classic Simulation Shows How Quickly the Risk Can Rise

Simmons, Nelson, and Simonsohn demonstrated the effect using simulations of common researcher degrees of freedom. In one scenario, researchers began with 20 observations per condition, tested for significance, and, if the result was nonsignificant, added 10 observations per condition before testing again.

That modest flexibility increased the simulated false-positive rate from the nominal 5% to 7.7%. When the authors examined more aggressive optional stopping, beginning with 10 observations per condition and testing after every additional observation per condition until reaching 50, the false-positive rate reached 22%.

These numbers are examples from particular simulations, not universal estimates for optional stopping. The amount of inflation depends on factors such as when testing begins, how frequently researchers look, the maximum sample size, the stopping criterion, and the statistical model.

Optional Stopping Is Not Simply “Changing Your Sample Size”

Sample sizes sometimes change for legitimate reasons.

Recruitment may be slower than expected. Funding may end. Equipment may fail. A blinded estimate of nuisance parameters may justify sample-size re-estimation under an appropriate design. Ethical or safety considerations can also require early termination.

Optional stopping refers more specifically to allowing accumulating information from the data to determine whether sampling continues.

The particularly familiar questionable version is significance-driven stopping:

If p <.05 Stop collecting data and report the finding.
If p ≥.05 Collect more observations and test again.

The asymmetry matters. A significant fluctuation ends the experiment, while a nonsignificant fluctuation receives another chance to change.

The Problem Is Not That the P-Value “Knows” You Peeked

It can be tempting to explain optional stopping by saying that looking at the data somehow contaminates them. It does not.

The issue concerns the long-run behavior of the decision procedure.

If you repeatedly simulate studies in which the null hypothesis is true and apply a rule of “continue after nonsignificance, stop after significance,” more than the nominal proportion of studies can eventually terminate with a significant result.

The final p-value may look exactly like one obtained from a fixed-sample study. What differs is the selection mechanism that determined when that final value was observed.

Stopping Rules Can Also Affect Which Effect Sizes Get Observed

Imagine that the true effect is zero. As observations accumulate, the estimated effect fluctuates around zero. Occasionally, random variation produces an unusually large estimate.

If you stop precisely when that fluctuation becomes large enough to produce statistical significance, unusually favorable estimates have a greater chance of becoming the final reported estimate.

This is one reason stopping procedures can affect not only hypothesis testing but also estimation and interpretation.

Optional Stopping Can Be One Form of P-Hacking

Optional stopping becomes especially concerning when it functions as one of several result-dependent researcher degrees of freedom.

A researcher might first add participants, then alter exclusions, try another covariate, inspect another outcome, and stop once some combination produces significance. Simmons and colleagues showed how combining several such freedoms can produce much larger false-positive rates than any single choice alone.

This is why p-hacking can emerge from ordinary-looking research decisions rather than from one obviously improper statistical maneuver.

Sequential Analysis Solves a Different Problem

Researchers often have legitimate reasons to examine accumulating data. In clinical research, for example, it may be ethically undesirable to continue exposing participants to an inferior treatment when strong evidence has already emerged. Researchers may also want efficient designs that can stop early for efficacy or futility.

Statistical sequential methods were developed precisely for settings in which more than one look at the data is part of the design.

Rather than repeatedly applying an unchanged.05 threshold as though every look were the only one, sequential procedures specify how evidence will be evaluated across interim analyses. Depending on the design, this may involve stopping boundaries, alpha-spending approaches, or other inferential machinery.

Unplanned optional stopping Researchers repeatedly inspect accumulating results and let those results determine whether to continue, without using an inferential procedure designed for that stopping rule.
Planned sequential analysis Repeated looks and possible early stopping are explicitly incorporated into the study design and statistical procedure.

This distinction is why the broader question of whether you can stop collecting data when a result becomes significant cannot be answered without knowing the design.

More Frequent Peeking Can Create More Opportunities

How often you examine the accumulating results matters.

Simmons and colleagues found in their simulations that false-positive rates increased as researchers were allowed to test more frequently while continuing sampling after nonsignificant results. Their most intensive simulated procedure produced substantially more false positives than a single additional sample-size decision.

This does not imply a universal formula such as “every peek adds X percentage points.” The tests are statistically dependent, and the operating characteristics depend on the entire procedure.

The general principle is simpler: if your decision rule repeatedly gives chance variation opportunities to trigger a favorable stopping condition, analyze the procedure as sequential rather than pretending the final look occurred in isolation.

Watch Out

Do not interpret the nominal p-value from the final ordinary test as automatically preserving its usual fixed-sample Type I error rate when that test was selected by repeatedly looking and stopping at significance. The inferential procedure includes the stopping rule, not just the final calculation.

04 · A Practical Example

How a Null Effect Can Eventually Produce a Significant Stopping Point

Hypothetical Example

Testing a Teaching Intervention With No Genuine Effect

Imagine, for illustration, that a teaching intervention truly has no effect on examination scores. A researcher nevertheless uses an optional stopping rule: begin with 40 students, test the group difference, and add 20 students whenever the result remains nonsignificant.

40 students: p =.31 Nothing significant appears, so data collection continues.
60 students: p =.14 Still nonsignificant. More students are added.
80 students: p =.07 The result looks promising, so the researcher continues.
100 students: p =.038 Random sampling variation finally carries the result below.05. The researcher stops.

This sequence is hypothetical, and another sequence of random samples could behave very differently. That is exactly the point.

If the researcher had committed to 140 participants, the result might later have moved above.05 again. Statistical significance at an intermediate sample size is not guaranteed to persist as observations accumulate.

Optional stopping systematically retains paths that happen to cross the boundary while allowing paths that have not crossed it to keep trying. Across repeated studies under the null, that asymmetry can produce more false-positive declarations than the nominal fixed-sample rate.

05 · What Researchers Often Get Wrong

Common Misunderstandings About Optional Stopping

Misconception

If P <.05, the Stopping Rule No Longer Matters

The stopping procedure helps determine the long-run error properties of the analysis. A final p-value below.05 does not erase the repeated opportunities through which that particular value was selected.

Misconception

One Extra Look Cannot Make Any Meaningful Difference

The effect of one additional look depends on the procedure, but it is not automatically zero. In Simmons and colleagues' simulation, allowing one additional sample-size increment after an initial nonsignificant test increased the false-positive rate from 5% to 7.7%.

Misconception

All Optional Stopping Produces the Same False-Positive Rate

No. The consequences depend on the minimum and maximum sample sizes, frequency of testing, stopping boundary, statistical model, and inferential framework. A particular simulation result should not be generalized into a universal inflation factor.

Misconception

Any Study That Stops Early Is Statistically Invalid

No. Sequential and adaptive methods can explicitly accommodate interim analyses and early stopping. The relevant distinction is whether the stopping procedure is incorporated appropriately into the statistical design.

Misconception

Optional Stopping Requires Dishonest Intent

No. A researcher may genuinely believe that collecting “just a few more” observations after a near-significant result is harmless. The statistical consequences arise from the procedure, not from whether the researcher intended to manufacture a false positive.

06 · What This Means for You

Specify How Sampling Will End Before the Results Become Tempting

A clear stopping rule removes one important source of result-dependent flexibility.

A simple decision framework

If your study uses a conventional fixed sample
Determine the sample-size or stopping criterion independently of whether interim results happen to be significant.
If you expect to examine accumulating results
Use an appropriate sequential or adaptive method in which interim analyses and stopping decisions are part of the statistical design.
If recruitment changes for an external practical reason
Document the reason and assess its implications rather than presenting the final sample size as though it had always been planned.
If you have already repeatedly tested the accumulating data
Preserve that information, disclose consequential deviations, and use statistical methods appropriate to the procedure that actually occurred where possible.
If your current result is “almost significant”
Do not let proximity to.05 become the rationale for adding observations unless your design already provides a valid rule for doing so.

The awkward moment often comes at p =.06 rather than p =.60. That is precisely when the stopping rule should already exist. Otherwise, methodological planning has a suspicious tendency to become much more creative when significance is only a few participants away.

07 · A Quick Checklist

Check Whether Your Stopping Rule Is Result-Dependent

Before or during data collection, check:
Have you specified when data collection will end independently of the desired result?
Will you inspect inferential results before reaching that stopping point?
If interim analyses are planned, does your statistical method explicitly accommodate them?
Would a significant interim result make you stop while a nonsignificant result would make you continue?
Have you specified the number and timing of interim looks where the design requires them?
If your sample size changed, have you documented why and when the decision occurred?
Are you reporting unplanned interim analyses rather than presenting the final sample as prospectively fixed?
Have you considered whether the stopping procedure affects effect estimation as well as false-positive control?
08 · Frequently Asked Questions

Frequently Asked Questions About Optional Stopping

Does optional stopping mean simply stopping before the planned sample size?

No. The key feature is that accumulating information influences the stopping decision. A study that ends early because a laboratory closes has a different stopping mechanism from one that ends because the p-value has just crossed.05.

Why does continuing after p >.05 increase false positives?

Because nonsignificant paths receive additional opportunities to cross the rejection threshold, while significant paths stop and are retained as successes. Under ordinary repeated unadjusted testing, this selection can increase the long-run probability of rejecting a true null hypothesis.

How much does optional stopping inflate the false-positive rate?

There is no single number. It depends on the stopping rule and testing procedure. Simmons and colleagues reported rates ranging from 7.7% for one modest sample-size flexibility scenario to substantially higher values under more frequent repeated testing. Those figures describe their simulated procedures, not all possible designs.

Is stopping at significance always p-hacking?

No. A properly designed sequential procedure can legitimately stop when a prespecified evidential boundary is reached. It becomes a p-hacking concern when researchers improvise result-dependent stopping while analyzing and reporting the study as though a conventional fixed-sample procedure had been followed.

Can I look at descriptive statistics while collecting data?

Whether interim inspection is consequential depends on what information you examine and whether it influences sampling or analysis decisions. Looking at accumulating outcome differences and allowing them to determine continuation is particularly relevant. High-stakes or complex designs should use a prospectively specified monitoring plan.

What if I need more participants because the data are noisier than expected?

Some designs allow sample-size adaptation based on appropriate information, including certain blinded nuisance-parameter estimates. The validity depends on how the decision is made and analyzed. “The p-value is too large, so we need more participants” is not equivalent to a statistically planned sample-size re-estimation procedure.

Does preregistering a maximum sample size solve optional stopping?

Not by itself. A maximum tells readers how large the study might become, but a sequential design also needs an appropriate rule governing interim looks and decisions. “Stop whenever p <.05, otherwise continue until N = 500” still describes result-dependent stopping.

09 · The Bottom Line

The Final P-Value Cannot Be Separated From How You Decided to Stop

The Bottom Line

Optional stopping can inflate false positives when researchers repeatedly apply ordinary significance tests to accumulating data, continue after nonsignificant results, and stop when significance appears, because the study has created multiple opportunities for random variation to trigger rejection.

The solution is not to prohibit all interim analysis. Decide prospectively whether your study is fixed-sample or sequential and use an inferential method appropriate to that design. A valid stopping rule is part of the statistical procedure, not an administrative detail added after the p-value becomes interesting.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes