01 · The Question
Why Does Checking the Data More Than Once Change the False-Positive Risk?
Suppose you begin with 40 participants and run your planned analysis. The result is not statistically significant, so you collect another 20 and test again. Still nothing. You add another 20, repeat the analysis, and continue until either p falls below.05 or you reach your maximum sample size.
You have not fabricated observations or altered the statistical test. Each individual analysis may look perfectly ordinary.
Yet the overall procedure is no longer equivalent to performing one fixed-sample test at the end. By repeatedly giving random variation opportunities to cross the significance threshold and stopping when it does, you can increase the probability of declaring an effect when no genuine effect exists.
03 · What You Need to Know
How Optional Stopping Changes the Statistical Experiment
The Stopping Rule Is Part of the Procedure
Consider two studies that both finish with 100 participants.
Study A planned from the beginning to collect 100 observations and perform one hypothesis test afterward.
Study B began with 40 observations, tested the hypothesis, and then repeatedly added participants whenever the result remained nonsignificant. It happened to reach significance when the sample reached 100, so collection stopped.
The final sample sizes are identical. The procedures that produced them are not.
Study B gave the null hypothesis several opportunities to be rejected. The decision to reach 100 observations depended partly on what happened at earlier sample sizes.
Why Repeated Opportunities Can Inflate False Positives
Under a conventional fixed-sample frequentist test conducted at an alpha level of.05, the testing procedure is designed so that, under the relevant assumptions and when the null hypothesis is true, the long-run probability of rejection is 5%.
Now imagine repeatedly applying an ordinary.05 test to accumulating data and stopping the first time significance appears.
At one interim look, random sampling variation may push the result below.05. If it does, the study stops and records a “positive” result. If it does not, the researcher gets another opportunity later.
Those repeated opportunities are correlated because later datasets contain earlier observations, so you cannot calculate the overall false-positive probability by simply multiplying the number of looks by 5%. Nevertheless, under ordinary unadjusted repeated testing, the overall probability can exceed 5%.
A Classic Simulation Shows How Quickly the Risk Can Rise
Simmons, Nelson, and Simonsohn demonstrated the effect using simulations of common researcher degrees of freedom. In one scenario, researchers began with 20 observations per condition, tested for significance, and, if the result was nonsignificant, added 10 observations per condition before testing again.
That modest flexibility increased the simulated false-positive rate from the nominal 5% to 7.7%. When the authors examined more aggressive optional stopping, beginning with 10 observations per condition and testing after every additional observation per condition until reaching 50, the false-positive rate reached 22%.
These numbers are examples from particular simulations, not universal estimates for optional stopping. The amount of inflation depends on factors such as when testing begins, how frequently researchers look, the maximum sample size, the stopping criterion, and the statistical model.
Optional Stopping Is Not Simply “Changing Your Sample Size”
Sample sizes sometimes change for legitimate reasons.
Recruitment may be slower than expected. Funding may end. Equipment may fail. A blinded estimate of nuisance parameters may justify sample-size re-estimation under an appropriate design. Ethical or safety considerations can also require early termination.
Optional stopping refers more specifically to allowing accumulating information from the data to determine whether sampling continues.
The particularly familiar questionable version is significance-driven stopping:
If p <.05
Stop collecting data and report the finding.
If p ≥.05
Collect more observations and test again.
The asymmetry matters. A significant fluctuation ends the experiment, while a nonsignificant fluctuation receives another chance to change.
The Problem Is Not That the P-Value “Knows” You Peeked
It can be tempting to explain optional stopping by saying that looking at the data somehow contaminates them. It does not.
The issue concerns the long-run behavior of the decision procedure.
If you repeatedly simulate studies in which the null hypothesis is true and apply a rule of “continue after nonsignificance, stop after significance,” more than the nominal proportion of studies can eventually terminate with a significant result.
The final p-value may look exactly like one obtained from a fixed-sample study. What differs is the selection mechanism that determined when that final value was observed.
Stopping Rules Can Also Affect Which Effect Sizes Get Observed
Imagine that the true effect is zero. As observations accumulate, the estimated effect fluctuates around zero. Occasionally, random variation produces an unusually large estimate.
If you stop precisely when that fluctuation becomes large enough to produce statistical significance, unusually favorable estimates have a greater chance of becoming the final reported estimate.
This is one reason stopping procedures can affect not only hypothesis testing but also estimation and interpretation.
Optional Stopping Can Be One Form of P-Hacking
Optional stopping becomes especially concerning when it functions as one of several result-dependent researcher degrees of freedom.
A researcher might first add participants, then alter exclusions, try another covariate, inspect another outcome, and stop once some combination produces significance. Simmons and colleagues showed how combining several such freedoms can produce much larger false-positive rates than any single choice alone.
This is why p-hacking can emerge from ordinary-looking research decisions rather than from one obviously improper statistical maneuver.
Sequential Analysis Solves a Different Problem
Researchers often have legitimate reasons to examine accumulating data. In clinical research, for example, it may be ethically undesirable to continue exposing participants to an inferior treatment when strong evidence has already emerged. Researchers may also want efficient designs that can stop early for efficacy or futility.
Statistical sequential methods were developed precisely for settings in which more than one look at the data is part of the design.
Rather than repeatedly applying an unchanged.05 threshold as though every look were the only one, sequential procedures specify how evidence will be evaluated across interim analyses. Depending on the design, this may involve stopping boundaries, alpha-spending approaches, or other inferential machinery.
Unplanned optional stopping
Researchers repeatedly inspect accumulating results and let those results determine whether to continue, without using an inferential procedure designed for that stopping rule.
Planned sequential analysis
Repeated looks and possible early stopping are explicitly incorporated into the study design and statistical procedure.
This distinction is why the broader question of whether you can stop collecting data when a result becomes significant cannot be answered without knowing the design.
More Frequent Peeking Can Create More Opportunities
How often you examine the accumulating results matters.
Simmons and colleagues found in their simulations that false-positive rates increased as researchers were allowed to test more frequently while continuing sampling after nonsignificant results. Their most intensive simulated procedure produced substantially more false positives than a single additional sample-size decision.
This does not imply a universal formula such as “every peek adds X percentage points.” The tests are statistically dependent, and the operating characteristics depend on the entire procedure.
The general principle is simpler: if your decision rule repeatedly gives chance variation opportunities to trigger a favorable stopping condition, analyze the procedure as sequential rather than pretending the final look occurred in isolation.
Watch Out
Do not interpret the nominal p-value from the final ordinary test as automatically preserving its usual fixed-sample Type I error rate when that test was selected by repeatedly looking and stopping at significance. The inferential procedure includes the stopping rule, not just the final calculation.
04 · A Practical Example
How a Null Effect Can Eventually Produce a Significant Stopping Point
Hypothetical Example
Testing a Teaching Intervention With No Genuine Effect
Imagine, for illustration, that a teaching intervention truly has no effect on examination scores. A researcher nevertheless uses an optional stopping rule: begin with 40 students, test the group difference, and add 20 students whenever the result remains nonsignificant.
40 students: p =.31
Nothing significant appears, so data collection continues.
60 students: p =.14
Still nonsignificant. More students are added.
80 students: p =.07
The result looks promising, so the researcher continues.
100 students: p =.038
Random sampling variation finally carries the result below.05. The researcher stops.
This sequence is hypothetical, and another sequence of random samples could behave very differently. That is exactly the point.
If the researcher had committed to 140 participants, the result might later have moved above.05 again. Statistical significance at an intermediate sample size is not guaranteed to persist as observations accumulate.
Optional stopping systematically retains paths that happen to cross the boundary while allowing paths that have not crossed it to keep trying. Across repeated studies under the null, that asymmetry can produce more false-positive declarations than the nominal fixed-sample rate.