01 · The Question
If You Try More Than One Analysis, Have You P-Hacked?
Your first analysis does not always have to be your last. Perhaps a model assumption fails. A reviewer requests another specification. You want to know whether an effect survives a different exclusion rule. Or you deliberately compare several plausible models because no single specification is unquestionably correct.
Then one analysis produces a statistically significant result while another does not. Have you crossed into p-hacking simply because you tried more than one?
No. Counting analyses is the wrong diagnostic. What matters is why the analyses were performed, how they were selected, what inferential claims are made from them, and what readers are ultimately shown .
03 · What You Need to Know
The Difference Between Analytical Investigation and Significance Hunting
Researchers Often Have More Than One Defensible Analysis
Statistical analysis is rarely a matter of pressing one uniquely correct button. Researchers may have legitimate choices about variable coding, functional form, missing-data procedures, covariates, exclusion criteria, estimators, distributional assumptions, or model structure.
Sometimes those alternatives are part of the scientific question. You might want to know whether a finding is robust to several defensible specifications. In other cases, diagnostic information reveals that the original model is inappropriate and should be revised.
The existence of researcher degrees of freedom therefore does not mean researchers should pretend that uncertainty about analysis does not exist. The more useful goal is to manage that flexibility without allowing favorable outcomes to determine which analytical path receives privileged status.
Purpose Matters: Are You Testing Robustness or Searching for Success?
Consider two researchers who each run six regression models.
Researcher A
Researcher B
Defines six theoretically defensible specifications to assess robustness
Keeps modifying the model after each nonsignificant result
Examines whether conclusions are consistent across them
Stops when one specification gives p <.05
Reports the consequential specifications and their differences
Reports only the successful specification
Interprets instability as uncertainty about the finding
Treats the successful result as decisive evidence
Both researchers ran multiple analyses. Their inferential procedures are not equivalent.
For Researcher A, variation across analyses is information. For Researcher B, variation is an obstacle to be searched until a desirable answer appears.
Robustness Analysis Asks Whether the Conclusion Survives Reasonable Choices
A sensitivity or robustness analysis intentionally examines what happens when reasonable analytical decisions change. For example, a researcher might estimate a relationship using alternative defensible measures, compare results with and without influential observations, or assess several plausible model specifications.
If the substantive conclusion remains similar, that stability may strengthen confidence that it is not an artifact of one arbitrary choice. If the conclusion changes substantially, that instability is itself scientifically important.
P-hacking reverses this logic. Rather than asking, “Does my conclusion survive reasonable alternatives?”, the implicit question becomes, “Which alternative gives me the conclusion I want?”
Robustness checking
Uses alternative reasonable analyses to learn whether conclusions depend on analytical choices.
Result-driven specification searching
Uses alternative analyses to locate a preferred result and privileges that result because it is favorable.
Exploratory Analysis Is Also Legitimate
Researchers are allowed to explore. Indeed, exploratory analysis can reveal unexpected patterns, generate hypotheses, identify nonlinear relationships, uncover heterogeneity, and expose problems with measurements or models.
The Center for Open Science explicitly recognizes exploratory and confirmatory analyses as valuable while emphasizing the importance of distinguishing between them. If you explore 30 relationships and discover an interesting association, that finding can be reported. What changes is the strength and type of inference you can reasonably make from the same data that generated the discovery.
You should therefore report exploratory findings as exploratory rather than reconstructing the analysis as though the successful relationship had been uniquely predicted beforehand.
Data-Dependent Decisions Can Matter Even If You Run Only One Final Test
There is a subtler complication. P-hacking and related problems cannot always be identified by asking how many analyses appear in your software history.
Gelman and Loken's “garden of forking paths” illustrates how an analysis can depend on observed data even when a researcher does not consciously run dozens of tests. Suppose you notice that an effect appears especially pronounced among older participants and, because that pattern looks theoretically meaningful, decide that age should define the focal subgroup. Had a different pattern appeared, you might have chosen a different analysis.
Only one focal test may ultimately be computed, but the path to that test was informed by the data. The set of potential analyses therefore matters, not merely the number of commands actually executed.
Watch Out
“I only ran one final test” does not necessarily establish that the test was selected independently of the data. Conversely, “I ran ten models” does not establish p-hacking. The analytical decision process matters more than the raw count.
Multiple Testing and P-Hacking Are Related but Not Identical
If you conduct many statistical tests, the probability of obtaining at least one apparently significant result by chance can increase, depending on the tests and inferential framework. This is a multiplicity problem even when every test is reported honestly.
P-hacking adds a selection problem: analytical choices are searched or adapted based on results, and favorable outcomes receive preferential treatment.
Researchers have several ways to address multiplicity depending on the design and purpose of the analysis. These may include limiting prespecified confirmatory tests, using appropriate multiplicity adjustments, modeling related effects jointly, emphasizing estimation and uncertainty rather than binary significance decisions, and clearly separating exploratory analyses from confirmatory claims.
The appropriate solution depends on the statistical problem. There is no universal instruction to “correct every p -value whenever more than one analysis exists.”
Changing an Analysis Can Be Methodologically Necessary
Suppose your prespecified linear model shows severe assumption problems. Keeping it simply because it was planned would not rescue the inference. You may need another model.
Similarly, you might discover a coding error, an invalid measurement, an unforeseen pattern of missingness, or an observation that should have been excluded under the study's substantive eligibility rules.
The appropriate response is to make the defensible change and document it. The National Academies recommends describing analytical decisions, decisions about data inclusion and exclusion, when consequential decisions were made, and whether analyses were exploratory or confirmatory.
Transparency should not become a methodological straitjacket. The goal is not obedience to an analysis plan after it becomes demonstrably unsuitable. It is an accurate account of how the final inference was reached.
Selective Reporting Is Often What Makes the Search Especially Misleading
Suppose you estimate eight reasonable models and they produce materially different results. Reporting only the strongest model deprives readers of information about the finding's dependence on specification.
This is why p-hacking frequently overlaps with selective reporting and other questionable research practices . The analytical search creates opportunities for a favorable result, while selective reporting makes that result look less contingent than it really is.
The American Statistical Association's Task Force on Statistical Significance and Replicability has similarly cautioned that model choice, insufficient description of analytical procedures, and selection of results to report can contribute to replicability problems. It also notes that even highlighting a few persuasive findings among a larger set may distort the apparent evidence.
04 · A Practical Example
The Same Six Analyses Can Tell Two Very Different Methodological Stories
Hypothetical Example
Does Online Discussion Improve Academic Performance?
A researcher estimates the association between participation in online discussions and final examination scores. Several defensible decisions exist about prior-achievement adjustment, attendance, and treatment of highly influential observations.
Scenario A: Robustness Checking
Define plausible specifications
The researcher identifies several substantively defensible models and explains what each tests.
Compare the results
The estimated association is positive in all six models, although uncertainty and effect estimates vary.
Report the variation
The researcher presents the principal analysis alongside consequential alternatives and explains the degree of robustness.
Interpret accordingly
The conclusion reflects the full pattern rather than whichever model has the smallest p -value.
Scenario B: Significance Hunting
Model 1: p =.14
The researcher adds another covariate.
Model 2: p =.09
A different exclusion rule is tried.
Models 3–5: still nonsignificant
Additional combinations are tested.
Model 6: p =.047
The search stops.
Only Model 6 is emphasized
The manuscript presents this specification as the analysis supporting the claimed association.
The number of analyses is identical. What differs is the role the observed results played in selecting and interpreting them.
In the first scenario, multiple analyses reveal how much the conclusion depends on modeling choices. In the second, multiple analyses function as repeated opportunities to obtain a preferred threshold-crossing result.
06 · What This Means for You
What to Do When Several Analyses Are Reasonable
Do not choose between “run only one analysis” and “run everything and keep the nicest result.” Neither is a satisfactory general strategy.
Instead, identify what each analysis is doing and what question the variation among analyses can answer.
A simple decision framework
If one analysis was clearly prespecified as the primary confirmatory test
Report that result regardless of whether it crosses a significance threshold, alongside justified deviations or supplementary analyses.
If several specifications are genuinely plausible
Examine whether the substantive conclusion is robust across them rather than choosing solely by statistical significance.
If diagnostics show the original analysis is inappropriate
Use a more appropriate analysis and explain why the change was necessary.
If an alternative was suggested by an interesting feature of the observed data
Treat the resulting inference as data-informed or exploratory where appropriate and report how the analysis arose.
If results conflict across reasonable analyses
Treat that instability as substantive information rather than hiding the specifications that weaken the preferred conclusion.
A useful warning sign is the phrase, “Let's try one more thing” repeated only after unfavorable results. Sometimes one more analysis really is necessary. But if the search reliably ends when p drops below.05, the stopping rule is telling you more about the analytical process than the methodological rationale is.
07 · A Quick Checklist
Before Choosing Among Multiple Analyses
When several analyses are possible, check:
What substantive or statistical reason justifies each alternative analysis?
Were any primary analyses or decision rules specified before you examined the relevant results?
Are you using alternative analyses to assess robustness or mainly to search for statistical significance?
Would you still choose the same specification if it produced a less favorable result?
Do reasonable specifications lead to materially different conclusions?
Have you considered multiplicity when several inferential tests address the same claim?
Are consequential deviations and data-dependent decisions clearly disclosed?
Does your manuscript distinguish prespecified confirmatory analyses from later exploratory analyses?
11 · Cite this Guide
How to Cite This Guide
This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.
Recommended (Field Guide)
APA
MLA
Chicago
Copy Citation