Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Can You Stop Collecting Data When the Result Becomes Statistically Significant?

In an ordinary fixed-sample design, repeatedly checking results and stopping when p < .05 can invalidate the usual interpretation of the test. Properly planned sequential designs are different.

494
Stopping Data Collection at Significance Guide 494 of 530
01 · The Question

Your Result Just Reached P <.05. Can You Stop Now?

You planned to collect 200 observations. After 120, curiosity wins and you run the analysis. The result is p =.08, so you continue. At 140 it becomes p =.06. At 160 it reaches p =.047.

Can you stop collecting data because the result is now statistically significant?

If your study was designed as an ordinary fixed-sample analysis with a conventional significance test, stopping because the observed p-value finally crossed the threshold changes the statistical procedure you are actually using. Repeated opportunities to stop for significance can increase the chance of a false-positive conclusion.

02 · The Short Answer

Do Not Turn P <.05 Into an Unplanned Stopping Rule

In Brief

In a conventional fixed-sample design, you generally should not repeatedly check the result and stop collecting data simply when statistical significance appears, because the usual test assumes a sampling and analysis plan that does not provide repeated unadjusted opportunities to declare significance.

Early stopping can be legitimate in properly designed sequential or adaptive studies with prospectively specified interim analyses and statistical procedures that account for them. The problem is not early stopping itself; it is using an unplanned significance-dependent stopping rule while interpreting the final result as though only one fixed-sample test had been performed.

03 · What You Need to Know

Why Repeatedly Checking for Significance Changes the Test

A Fixed-Sample Test Assumes More Than the Final Sample Size

Suppose you design a study for 200 participants and plan one hypothesis test after all 200 observations have been collected.

That procedure is different from collecting up to 200 participants while testing after every additional group of observations and stopping whenever p falls below.05.

The final dataset might contain 160 participants in either case. But the sampling procedures are different.

In the second procedure, the null hypothesis has had repeated opportunities to be rejected. A temporary fluctuation can trigger stopping before later observations have an opportunity to move the estimate back toward the null.

Why the False-Positive Probability Can Increase

A conventional significance threshold such as.05 is calibrated for a particular testing procedure. It does not mean that researchers may look repeatedly at accumulating data, perform the same ordinary test each time at.05, and still automatically retain an overall 5% Type I error probability.

The FDA's guidance on adaptive clinical trials makes the point explicitly for group sequential designs: performing each of multiple efficacy tests at the conventional significance level inflates the Type I error probability. Sequential designs therefore use methods that determine appropriate stopping boundaries so that the overall error probability is controlled.

The exact inflation depends on the stopping procedure, number and timing of looks, dependence among tests, and statistical method. There is no single universal inflation factor that applies to every form of repeated testing.

This Practice Is Commonly Called Optional Stopping

When the decision to continue or stop sampling depends on the accumulating results, the procedure is often described as optional stopping.

One simple version looks like this:

Check 1 The result is nonsignificant, so collect more observations.
Check 2 The result is still nonsignificant, so continue again.
Check 3 The result reaches p <.05, so stop immediately.

The stopping decision is now coupled to the inferential result. The study does not simply have a sample size; it has a data-dependent sampling rule.

The statistical consequences and the conditions under which they arise are examined more closely in the discussion of optional stopping and false positives.

Stopping Early Is Not Inherently Invalid

There are well-established statistical designs in which researchers examine accumulating data and may legitimately stop before reaching the maximum sample size.

Group sequential designs, for example, allow one or more prospectively planned interim analyses with specified criteria for stopping. FDA guidance notes that these designs can permit early stopping for efficacy or futility while using stopping boundaries designed to control Type I error.

Adaptive designs can go further, allowing prospectively planned modifications based on accumulating information without undermining trial validity when designed and implemented appropriately. FDA guidance defines adaptive designs in terms of prospectively planned modifications and emphasizes evaluation of their operating characteristics.

Unplanned significance stopping Researchers repeatedly apply an ordinary test and decide after seeing the result whether to continue, stopping when the desired threshold appears.
Planned sequential or adaptive design Interim looks, decision rules, maximum sample size, and appropriate inferential procedures are prospectively designed so that repeated examination is part of the statistical method.

“I Only Peeked Once” Still Changes the Planned Procedure

Suppose you planned one analysis at 200 participants but check at 150. The result is significant, so you stop.

This is less extensive than checking every ten observations, but the decision to stop still depended on an interim result that was not part of the original fixed-sample procedure.

The statistical consequences may be smaller than under aggressive repeated peeking, but “only once” does not make the additional opportunity nonexistent.

If interim analysis may be needed, it is better to plan it prospectively and use a method designed for that purpose.

Stopping Because the Result Is Nonsignificant Can Also Be Data-Dependent

Optional stopping is not limited to stopping for success.

Researchers might stop because the result looks hopeless, because the estimated effect is smaller than expected, or because further data appear unlikely to change the conclusion.

Formal sequential designs can include futility rules. But an informal decision to stop after looking at disappointing results is still data-dependent and may have consequences for inference, bias, and interpretation.

Not every practical interruption in recruitment creates the same problem. A study might stop because funding expires, equipment fails, recruitment becomes impossible, or an external event makes continuation unethical. The reason for stopping matters and should be reported when it affects interpretation.

Continuing After Significance Can Be Result-Dependent Too

The mirror image deserves attention.

Suppose your result becomes significant, but the effect is smaller than hoped, so you continue collecting data in the hope that it becomes stronger. Or the result supports an unexpected direction, so you continue until it moves toward the original hypothesis.

The broader principle is that researchers should not allow the attractiveness of the accumulating result to determine sample size while subsequently analyzing the study as though sample size had been fixed independently.

Sample-Size Planning Is Not a Promise That Nothing Can Ever Change

Real studies encounter recruitment difficulties, attrition, variance estimates different from expectations, event rates that diverge from assumptions, and other complications.

Some designs permit blinded sample-size re-estimation or other planned adaptations. The relevant statistical methodology depends on the study and the information used to make the adaptation.

The lesson is not “never change sample size.” It is that a sample-size change should have a defensible design-based rationale and appropriate statistical treatment rather than being improvised according to whether the current p-value is attractive.

Watch Out

A maximum sample size does not by itself make repeated significance testing valid. “We planned to collect up to 300 participants” is not the same as prospectively specifying how and when interim analyses will occur and how those repeated looks will be handled statistically.

Early Stopping Can Also Affect Effect Estimates

Type I error is not the only issue.

When a study stops because an interim effect happens to look especially favorable, the observed effect at stopping can be exaggerated. FDA guidance for adaptive medical-device studies notes that even when Type I error is properly controlled, treatment-effect estimates after early stopping for success can be biased upward and may require prospectively planned methods for addressing that bias.

So a statistically valid early-stopping design still requires careful estimation and interpretation. Controlling false positives does not automatically solve every consequence of stopping early.

04 · A Practical Example

How “Just Checking” Can Become the Actual Stopping Rule

Hypothetical Example

A Study Planned for 200 Participants

A researcher plans to test whether a new instructional intervention improves assessment scores. The original design specifies 200 participants and one analysis after data collection is complete.

100 participants: p =.12 The researcher checks early. The result is not significant, so recruitment continues.
120 participants: p =.09 Another analysis is run. Recruitment continues.
140 participants: p =.054 The result is close, so another 20 participants are recruited.
160 participants: p =.041 The researcher stops because the result has crossed.05.
The manuscript The final analysis is reported as p =.041 without explaining that several interim tests determined when sampling stopped.

The problem is not that 160 participants is inherently an unacceptable sample size. Nor is p =.041 automatically false.

The problem is that the study used repeated significance testing to determine its sample size while the final analysis is interpreted as though the researchers had planned one ordinary test at 160 observations.

A legitimate sequential study could also stop at 160 participants. But its interim analyses and stopping boundaries would be part of the prospective statistical design rather than an improvised response to the accumulating p-value.

05 · What Researchers Often Get Wrong

Common Misunderstandings About Stopping Data Collection

Misconception

Once P <.05, Collecting More Data Is Unnecessary

Not in an ordinary fixed-sample study. The stopping rule is part of the statistical procedure. If your design calls for a predetermined sample size, an early threshold crossing does not automatically authorize stopping while retaining the same inferential interpretation.

Misconception

You Can Keep Collecting Data as Long as You Have Not Reached Your Maximum

A maximum alone does not specify a valid sequential design. If repeated interim results determine whether sampling continues, those repeated looks need to be incorporated into the statistical procedure where the inferential framework requires it.

Misconception

All Early Stopping Is P-Hacking

No. Group sequential and adaptive designs can legitimately stop early for efficacy, futility, safety, or other prospectively specified reasons. FDA guidance explicitly recognizes such designs when their rules and operating characteristics are appropriately planned.

Misconception

Stopping Early Only Affects the P-Value

No. Early stopping can also affect estimation. In some adaptive designs, stopping for an unusually favorable interim result can lead naïve treatment-effect estimates to be biased upward even when Type I error is controlled.

Misconception

If You Do Not Tell Anyone About the Interim Checks, the Final Test Is Still Valid

Reporting does not determine the statistical procedure that actually occurred. Concealing interim looks makes the study less transparent; it does not restore the operating characteristics of a test whose sampling rule depended on those looks.

06 · What This Means for You

Decide How You Will Stop Before the Result Makes the Decision for You

Before collecting data, determine whether your design is genuinely fixed-sample or whether interim decisions may be needed.

A simple decision framework

If you plan one analysis after a fixed sample is collected
Do not repeatedly test the hypothesis and stop merely because significance appears early.
If interim analyses are scientifically, ethically, or practically desirable
Use an appropriate sequential or adaptive design with prospectively specified decision rules and statistical methods.
If you have already conducted unplanned interim significance tests
Do not pretend they did not occur. Seek appropriate statistical advice about the implications for analysis and report consequential deviations transparently.
If recruitment must stop for an external reason
Document the actual reason and consider how the reduced sample size affects precision, power, and interpretation.
If the interim result looks overwhelmingly favorable
Follow the prespecified monitoring and stopping procedure rather than inventing a new rule because the result looks persuasive.

The correct sequential method depends on the design and inferential framework. Group sequential boundaries, alpha-spending approaches, adaptive designs, and Bayesian sequential procedures operate differently. They should not be collapsed into the folk rule “check whenever you like and stop when the result looks good.”

07 · A Quick Checklist

Before Looking at Results While Data Collection Is Still Running

Check your stopping procedure:
Was the intended sample size or stopping rule specified before examining the relevant results?
Does your design permit interim analysis, or was one final analysis originally planned?
Are the number and timing of interim looks specified appropriately for the design?
Does the statistical method account for repeated opportunities to stop or declare success?
Would you continue collecting data if the current p-value were nonsignificant but stop if it were significant?
If the sample size changed, is the reason methodological or operational rather than simply result-driven?
Have unplanned interim analyses or changes to the stopping rule been documented?
Have you considered how early stopping affects effect estimation as well as hypothesis testing?
08 · Frequently Asked Questions

Frequently Asked Questions About Stopping at Statistical Significance

Can I check my p-value before reaching the planned sample size?

You can technically calculate it, but if the result influences whether sampling continues, you have introduced an interim decision into the design. Conventional fixed-sample inference may no longer have the operating characteristics you assumed. If interim looks are needed, plan an appropriate sequential procedure.

What if the result is extremely significant?

A striking interim result does not automatically override the study design. Proper sequential designs can include boundaries for stopping when evidence becomes sufficiently strong, but those rules should be incorporated into the statistical procedure rather than invented after seeing the result.

Can I stop because continuing would waste resources?

Efficiency is a legitimate reason to consider sequential or adaptive designs prospectively. If an unforeseen operational constraint forces early termination, report it honestly and assess its implications. What is problematic is using ordinary significance itself as an unplanned resource-allocation rule while retaining fixed-sample inference.

Can I stop early if the study is clearly not working?

Formal designs can include futility stopping rules. An unplanned decision based on disappointing interim results is different and can affect interpretation. In high-stakes trials, stopping decisions may also involve independent monitoring and ethical considerations beyond statistical significance.

Is checking confidence intervals instead of p-values a solution?

Not automatically. If repeated inspection of any inferential statistic determines when you stop, the sequential nature of the procedure still matters. Valid sequential methods can be formulated using different statistical quantities; simply replacing a p-value with a confidence interval does not remove the stopping issue.

Does optional stopping always inflate false positives?

No universal statement applies to every inferential framework and stopping rule. Under ordinary repeated unadjusted frequentist significance testing, significance-dependent optional stopping can inflate Type I error. Properly constructed sequential, adaptive, or other inferential procedures can explicitly accommodate data-dependent stopping.

What if I already stopped when the result became significant?

Do not conceal the stopping process. Document how often the data were examined, how the stopping decision was made, and seek statistical guidance about an analysis appropriate to the procedure that actually occurred. Transparency cannot undo the design change, but hiding it compounds the problem.

09 · The Bottom Line

Early Stopping Can Be Valid, but “Stop When P <.05” Is Not a Free Option

The Bottom Line

If you planned an ordinary fixed-sample analysis, repeatedly checking the accumulating data and stopping as soon as statistical significance appears can inflate false-positive risk and invalidate the usual interpretation of the final test.

That does not mean researchers must always collect a fixed number of observations regardless of circumstances. Sequential and adaptive designs can legitimately permit early stopping, but the interim analyses and decision rules need to be part of an appropriate statistical design rather than improvised after the p-value starts looking attractive.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes