01 · The Question
Your Result Just Reached P <.05. Can You Stop Now?
You planned to collect 200 observations. After 120, curiosity wins and you run the analysis. The result is p =.08, so you continue. At 140 it becomes p =.06. At 160 it reaches p =.047.
Can you stop collecting data because the result is now statistically significant?
If your study was designed as an ordinary fixed-sample analysis with a conventional significance test, stopping because the observed p-value finally crossed the threshold changes the statistical procedure you are actually using. Repeated opportunities to stop for significance can increase the chance of a false-positive conclusion.
03 · What You Need to Know
Why Repeatedly Checking for Significance Changes the Test
A Fixed-Sample Test Assumes More Than the Final Sample Size
Suppose you design a study for 200 participants and plan one hypothesis test after all 200 observations have been collected.
That procedure is different from collecting up to 200 participants while testing after every additional group of observations and stopping whenever p falls below.05.
The final dataset might contain 160 participants in either case. But the sampling procedures are different.
In the second procedure, the null hypothesis has had repeated opportunities to be rejected. A temporary fluctuation can trigger stopping before later observations have an opportunity to move the estimate back toward the null.
Why the False-Positive Probability Can Increase
A conventional significance threshold such as.05 is calibrated for a particular testing procedure. It does not mean that researchers may look repeatedly at accumulating data, perform the same ordinary test each time at.05, and still automatically retain an overall 5% Type I error probability.
The FDA's guidance on adaptive clinical trials makes the point explicitly for group sequential designs: performing each of multiple efficacy tests at the conventional significance level inflates the Type I error probability. Sequential designs therefore use methods that determine appropriate stopping boundaries so that the overall error probability is controlled.
The exact inflation depends on the stopping procedure, number and timing of looks, dependence among tests, and statistical method. There is no single universal inflation factor that applies to every form of repeated testing.
This Practice Is Commonly Called Optional Stopping
When the decision to continue or stop sampling depends on the accumulating results, the procedure is often described as optional stopping.
One simple version looks like this:
Check 1
The result is nonsignificant, so collect more observations.
Check 2
The result is still nonsignificant, so continue again.
Check 3
The result reaches p <.05, so stop immediately.
The stopping decision is now coupled to the inferential result. The study does not simply have a sample size; it has a data-dependent sampling rule.
The statistical consequences and the conditions under which they arise are examined more closely in the discussion of optional stopping and false positives.
Stopping Early Is Not Inherently Invalid
There are well-established statistical designs in which researchers examine accumulating data and may legitimately stop before reaching the maximum sample size.
Group sequential designs, for example, allow one or more prospectively planned interim analyses with specified criteria for stopping. FDA guidance notes that these designs can permit early stopping for efficacy or futility while using stopping boundaries designed to control Type I error.
Adaptive designs can go further, allowing prospectively planned modifications based on accumulating information without undermining trial validity when designed and implemented appropriately. FDA guidance defines adaptive designs in terms of prospectively planned modifications and emphasizes evaluation of their operating characteristics.
Unplanned significance stopping
Researchers repeatedly apply an ordinary test and decide after seeing the result whether to continue, stopping when the desired threshold appears.
Planned sequential or adaptive design
Interim looks, decision rules, maximum sample size, and appropriate inferential procedures are prospectively designed so that repeated examination is part of the statistical method.
“I Only Peeked Once” Still Changes the Planned Procedure
Suppose you planned one analysis at 200 participants but check at 150. The result is significant, so you stop.
This is less extensive than checking every ten observations, but the decision to stop still depended on an interim result that was not part of the original fixed-sample procedure.
The statistical consequences may be smaller than under aggressive repeated peeking, but “only once” does not make the additional opportunity nonexistent.
If interim analysis may be needed, it is better to plan it prospectively and use a method designed for that purpose.
Stopping Because the Result Is Nonsignificant Can Also Be Data-Dependent
Optional stopping is not limited to stopping for success.
Researchers might stop because the result looks hopeless, because the estimated effect is smaller than expected, or because further data appear unlikely to change the conclusion.
Formal sequential designs can include futility rules. But an informal decision to stop after looking at disappointing results is still data-dependent and may have consequences for inference, bias, and interpretation.
Not every practical interruption in recruitment creates the same problem. A study might stop because funding expires, equipment fails, recruitment becomes impossible, or an external event makes continuation unethical. The reason for stopping matters and should be reported when it affects interpretation.
Continuing After Significance Can Be Result-Dependent Too
The mirror image deserves attention.
Suppose your result becomes significant, but the effect is smaller than hoped, so you continue collecting data in the hope that it becomes stronger. Or the result supports an unexpected direction, so you continue until it moves toward the original hypothesis.
The broader principle is that researchers should not allow the attractiveness of the accumulating result to determine sample size while subsequently analyzing the study as though sample size had been fixed independently.
Sample-Size Planning Is Not a Promise That Nothing Can Ever Change
Real studies encounter recruitment difficulties, attrition, variance estimates different from expectations, event rates that diverge from assumptions, and other complications.
Some designs permit blinded sample-size re-estimation or other planned adaptations. The relevant statistical methodology depends on the study and the information used to make the adaptation.
The lesson is not “never change sample size.” It is that a sample-size change should have a defensible design-based rationale and appropriate statistical treatment rather than being improvised according to whether the current p-value is attractive.
Watch Out
A maximum sample size does not by itself make repeated significance testing valid. “We planned to collect up to 300 participants” is not the same as prospectively specifying how and when interim analyses will occur and how those repeated looks will be handled statistically.
Early Stopping Can Also Affect Effect Estimates
Type I error is not the only issue.
When a study stops because an interim effect happens to look especially favorable, the observed effect at stopping can be exaggerated. FDA guidance for adaptive medical-device studies notes that even when Type I error is properly controlled, treatment-effect estimates after early stopping for success can be biased upward and may require prospectively planned methods for addressing that bias.
So a statistically valid early-stopping design still requires careful estimation and interpretation. Controlling false positives does not automatically solve every consequence of stopping early.
04 · A Practical Example
How “Just Checking” Can Become the Actual Stopping Rule
Hypothetical Example
A Study Planned for 200 Participants
A researcher plans to test whether a new instructional intervention improves assessment scores. The original design specifies 200 participants and one analysis after data collection is complete.
100 participants: p =.12
The researcher checks early. The result is not significant, so recruitment continues.
120 participants: p =.09
Another analysis is run. Recruitment continues.
140 participants: p =.054
The result is close, so another 20 participants are recruited.
160 participants: p =.041
The researcher stops because the result has crossed.05.
The manuscript
The final analysis is reported as p =.041 without explaining that several interim tests determined when sampling stopped.
The problem is not that 160 participants is inherently an unacceptable sample size. Nor is p =.041 automatically false.
The problem is that the study used repeated significance testing to determine its sample size while the final analysis is interpreted as though the researchers had planned one ordinary test at 160 observations.
A legitimate sequential study could also stop at 160 participants. But its interim analyses and stopping boundaries would be part of the prospective statistical design rather than an improvised response to the accumulating p-value.
06 · What This Means for You
Decide How You Will Stop Before the Result Makes the Decision for You
Before collecting data, determine whether your design is genuinely fixed-sample or whether interim decisions may be needed.
A simple decision framework
If you plan one analysis after a fixed sample is collected
Do not repeatedly test the hypothesis and stop merely because significance appears early.
If interim analyses are scientifically, ethically, or practically desirable
Use an appropriate sequential or adaptive design with prospectively specified decision rules and statistical methods.
If you have already conducted unplanned interim significance tests
Do not pretend they did not occur. Seek appropriate statistical advice about the implications for analysis and report consequential deviations transparently.
If recruitment must stop for an external reason
Document the actual reason and consider how the reduced sample size affects precision, power, and interpretation.
If the interim result looks overwhelmingly favorable
Follow the prespecified monitoring and stopping procedure rather than inventing a new rule because the result looks persuasive.
The correct sequential method depends on the design and inferential framework. Group sequential boundaries, alpha-spending approaches, adaptive designs, and Bayesian sequential procedures operate differently. They should not be collapsed into the folk rule “check whenever you like and stop when the result looks good.”
07 · A Quick Checklist
Before Looking at Results While Data Collection Is Still Running
Check your stopping procedure:
Was the intended sample size or stopping rule specified before examining the relevant results?
Does your design permit interim analysis, or was one final analysis originally planned?
Are the number and timing of interim looks specified appropriately for the design?
Does the statistical method account for repeated opportunities to stop or declare success?
Would you continue collecting data if the current p-value were nonsignificant but stop if it were significant?
If the sample size changed, is the reason methodological or operational rather than simply result-driven?
Have unplanned interim analyses or changes to the stopping rule been documented?
Have you considered how early stopping affects effect estimation as well as hypothesis testing?