01 · The Question
What Does P-Hacking Look Like Beyond the Definition?
“P-hacking” can sound like an obvious act: a researcher keeps manipulating an analysis until a statistically significant result appears. Real cases are often less theatrical.
A researcher might measure several outcomes and focus on the one with p <.05. Another might add a covariate after the original analysis misses significance. Someone else might collect a few more participants, rerun the test, and stop as soon as the threshold is crossed. Each individual decision can look reasonable when considered alone.
That is what makes p-hacking difficult to recognize. The warning sign is usually not a particular statistical technique. It is a result-dependent sequence of choices that gives favorable findings more opportunities to emerge and then hides or understates that analytical search.
03 · What You Need to Know
Recognizing P-Hacking in Ordinary Research Decisions
Testing Several Outcomes and Keeping the One That “Works”
Suppose a study measures anxiety, stress, sleep quality, life satisfaction, and academic engagement. The researcher tests the intervention against all five outcomes. Four comparisons are nonsignificant, while academic engagement produces p =.03.
There is nothing inherently wrong with measuring or analyzing several outcomes. The problem arises if the researcher presents academic engagement as though it were the study's sole or prespecified outcome while the other tests disappear from the research account.
The researcher has effectively given the data several opportunities to produce a result below the chosen significance threshold. Selecting the successful test afterward while interpreting its p -value as though only one test had been contemplated can overstate the evidence.
Trying Covariates Until Statistical Significance Appears
Covariates can improve an analysis when they are theoretically and statistically justified. They can also create substantial analytical flexibility.
Imagine that the unadjusted treatment effect produces p =.09. Controlling for age gives p =.07. Controlling for prior achievement gives p =.06. Including age, prior achievement, and socioeconomic status produces p =.04, so that specification becomes the one reported in the manuscript.
The problem is not that covariates were used. It is that statistical significance determined which specification became the apparent analysis. Simmons, Nelson, and Simonsohn identified flexibility involving covariates as one of several researcher degrees of freedom capable of increasing false-positive rates.
Changing Which Participants Count
Researchers often have legitimate reasons to exclude observations. A participant may fail an attention check, a sensor may malfunction, a response time may be physically implausible, or eligibility criteria may have been violated.
Now consider a different process. The full sample gives p =.08. Removing participants who completed unusually quickly gives p =.06. Removing another category of “outliers” produces p =.04, and only this final exclusion rule appears in the paper.
The exclusion itself might even sound defensible. What matters is whether the rule was selected because of its effect on the result. When several plausible inclusion rules are available, outcome-dependent selection among them can create a questionable use of analytical flexibility .
Searching Across Subgroups Until an Effect Appears
A treatment may show no clear effect in the full sample. Researchers then examine men and women separately, younger and older participants, high- and low-performing participants, several regions, different diagnostic groups, or combinations of these characteristics.
Eventually, one subgroup produces p <.05.
A scientifically motivated subgroup analysis can be perfectly legitimate. Searching many subgroups and presenting only the successful one as compelling evidence is different. As the number of investigated possibilities grows, so does the opportunity to encounter an apparently interesting pattern by chance.
The distinction becomes especially important when deciding when subgroup analysis turns into cherry-picking .
Changing How a Variable Is Defined
Many constructs can be operationalized in more than one defensible way. Achievement might be represented by a final examination score, course grade, standardized score, or composite measure. Age could be continuous or categorized. A scale might use all items or exclude an item with poor psychometric performance.
Alternative operationalizations can be useful for assessing robustness. They become problematic when researchers try several definitions, discover which produces the preferred inferential result, and report that definition as though it were the obvious choice all along.
Trying Different Statistical Models Until One Crosses the Threshold
The same research question can sometimes be analyzed using several plausible model specifications. Researchers might test different interaction terms, transformations, random-effects structures, estimators, or sets of predictors.
Again, model comparison is not itself p-hacking. Statistical modeling often requires examining alternatives. The danger lies in treating p <.05 as the model-selection criterion and then presenting the winning specification without revealing the search.
This is why trying several analyses until one works requires more nuance than simply counting how many models were run.
Collecting More Data After Checking the P-Value
Suppose a researcher initially collects 100 observations and obtains p =.08. Rather than following a predetermined stopping rule, the researcher adds 20 participants and checks again. The result becomes p =.06. Another 20 are collected, producing p =.04, and data collection stops.
This practice is commonly called optional stopping . Under conventional fixed-sample testing, repeatedly looking at the accumulating data and allowing the observed result to determine whether collection continues can alter the error properties of the procedure.
This does not mean researchers can never analyze accumulating data. Sequential and adaptive designs can formally permit interim analyses when the stopping procedure and statistical method appropriately account for them. The problem is unplanned peeking coupled with a significance-dependent decision to stop.
The Researcher May Not Feel Like They Are “Hacking” Anything
The term p-hacking can imply deliberate manipulation, but the underlying statistical problem does not require conscious dishonesty. Gelman and Loken described a related “garden of forking paths” problem in which analytical decisions are influenced by features of the observed data even when the researcher does not literally run every possible analysis.
A researcher might sincerely believe that an exclusion became appropriate after inspecting the data, that a subgroup suddenly became theoretically interesting, or that one operationalization now seems more sensible. Those judgments may even be reasonable. Yet if the observed data influence which inferential path is taken, the conventional interpretation of the final test may no longer correspond neatly to a single analysis selected independently of the data.
Watch Out
Intent is not a statistical correction. A researcher can introduce result-dependent flexibility without deliberately trying to deceive anyone. The relevant question is how the analytical procedure behaved, not merely whether the researcher felt they were “fishing.”
04 · A Practical Example
How a Nonsignificant Result Can Gradually Become “Significant”
Hypothetical Example
A Digital Learning Intervention
A researcher tests whether a digital learning intervention improves examination performance. The original plan is to compare intervention and control groups after 120 students have completed the study.
Planned analysis: p =.11
The primary comparison does not reach the conventional.05 threshold.
Add a covariate: p =.08
The researcher adjusts for students' previous grades.
Exclude several low-engagement participants: p =.06
A plausible engagement criterion is introduced after inspecting the data.
Analyze only first-year students: p =.03
The effect becomes statistically significant in one subgroup.
Report the successful analysis
The manuscript highlights the first-year subgroup and does not make the preceding analytical search clear.
No single step necessarily proves p-hacking. There could be substantive reasons for covariate adjustment, exclusions, and subgroup analysis. The problem becomes apparent from the selection mechanism : each new choice is evaluated partly by whether it moves the result toward statistical significance, and the eventual successful result is detached from the unsuccessful paths that led to it.
A more transparent analysis could report the planned result first, identify subsequent analyses as exploratory or sensitivity analyses, explain why they were performed, and show how conclusions change across reasonable specifications. The reader can then see that the apparent effect depends on particular analytical decisions.
06 · What This Means for You
How to Tell Whether Your Own Analysis Is Drifting Toward P-Hacking
When you change an analysis, ask what prompted the change. Would you have made the same decision if the previous result had already supported your hypothesis?
That counterfactual question is surprisingly useful. If you would exclude an observation only when including it produces p =.06, but happily retain it when the result is p =.03, statistical significance is influencing a methodological decision that should ordinarily have another justification.
A simple decision framework
If the alternative analysis tests robustness
Show whether the substantive conclusion changes across reasonable specifications rather than reporting only the preferred version.
If the data reveal a genuine methodological problem
Correct the analysis, document the reason, and disclose consequential departures from the original plan.
If you are exploring an unexpected pattern
Explore it, but identify the analysis as data-informed or exploratory rather than presenting it as a prespecified confirmatory test.
If you keep changing decisions because the result is nonsignificant
Stop treating statistical significance as the criterion for analytical choice. Reconsider the analysis on methodological grounds and disclose the search already conducted.
Preregistration or another prospective analysis plan can reduce ambiguity by documenting important decisions before the relevant results are known. It does not prohibit later exploration. The Center for Open Science explicitly distinguishes planned analyses from additional exploratory work and encourages transparent reporting of both.
07 · A Quick Checklist
Check Whether Your Analysis Shows Signs of P-Hacking
Before treating a result as confirmatory, check:
Did statistical significance influence which outcome you chose to emphasize?
Did you try alternative covariates, exclusions, transformations, or models and retain the version producing the preferred result?
Did you search multiple subgroups and report only the subgroup showing an effect?
Did you repeatedly check significance while collecting data and allow the result to determine when you stopped?
Would you have made the same analytical decisions if the earlier analysis had already supported your hypothesis?
Have you disclosed consequential analytical alternatives and deviations from your original plan?
Have you distinguished exploratory analyses from genuinely prespecified confirmatory analyses?
Does your interpretation reflect the full analytical search rather than only the most favorable result?
09 · The Bottom Line
P-Hacking Is Usually About the Process, Not One Statistical Test
The Bottom Line
P-hacking occurs when flexibility in collecting, analyzing, or selecting data is used in a result-dependent search for favorable statistical evidence and that search is not appropriately reflected in the analysis or reporting.
Changing a model, removing an observation, testing a subgroup, adding a covariate, or conducting another analysis is not automatically improper. Look at the sequence of decisions: what motivated each choice, how many plausible paths existed, what determined which result was emphasized, and whether readers are shown enough of that process to interpret the evidence accurately.
11 · Cite this Guide
How to Cite This Guide
This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.
Recommended (Field Guide)
APA
MLA
Chicago
Copy Citation