Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Does P-Hacking Actually Look Like in Real Research?

P-hacking is not limited to obviously manipulating data. It can emerge through ordinary-looking analytical choices when researchers use the results to decide what to test, change, or report.

485
What P-Hacking Looks Like Guide 485 of 530
01 · The Question

What Does P-Hacking Look Like Beyond the Definition?

“P-hacking” can sound like an obvious act: a researcher keeps manipulating an analysis until a statistically significant result appears. Real cases are often less theatrical.

A researcher might measure several outcomes and focus on the one with p <.05. Another might add a covariate after the original analysis misses significance. Someone else might collect a few more participants, rerun the test, and stop as soon as the threshold is crossed. Each individual decision can look reasonable when considered alone.

That is what makes p-hacking difficult to recognize. The warning sign is usually not a particular statistical technique. It is a result-dependent sequence of choices that gives favorable findings more opportunities to emerge and then hides or understates that analytical search.

02 · The Short Answer

What Are Some Concrete Examples of P-Hacking?

In Brief

P-hacking can include testing several outcomes and reporting the significant one, repeatedly changing covariates or exclusion rules, searching subgroups for an effect, trying alternative variable definitions, or repeatedly checking the data and stopping collection when statistical significance appears.

None of these actions is automatically p-hacking in isolation. The critical issue is whether analytical or data-collection choices are driven by obtaining a favorable result and then reported in a way that conceals the resulting multiplicity or flexibility.

03 · What You Need to Know

Recognizing P-Hacking in Ordinary Research Decisions

Testing Several Outcomes and Keeping the One That “Works”

Suppose a study measures anxiety, stress, sleep quality, life satisfaction, and academic engagement. The researcher tests the intervention against all five outcomes. Four comparisons are nonsignificant, while academic engagement produces p =.03.

There is nothing inherently wrong with measuring or analyzing several outcomes. The problem arises if the researcher presents academic engagement as though it were the study's sole or prespecified outcome while the other tests disappear from the research account.

The researcher has effectively given the data several opportunities to produce a result below the chosen significance threshold. Selecting the successful test afterward while interpreting its p-value as though only one test had been contemplated can overstate the evidence.

Trying Covariates Until Statistical Significance Appears

Covariates can improve an analysis when they are theoretically and statistically justified. They can also create substantial analytical flexibility.

Imagine that the unadjusted treatment effect produces p =.09. Controlling for age gives p =.07. Controlling for prior achievement gives p =.06. Including age, prior achievement, and socioeconomic status produces p =.04, so that specification becomes the one reported in the manuscript.

The problem is not that covariates were used. It is that statistical significance determined which specification became the apparent analysis. Simmons, Nelson, and Simonsohn identified flexibility involving covariates as one of several researcher degrees of freedom capable of increasing false-positive rates.

Changing Which Participants Count

Researchers often have legitimate reasons to exclude observations. A participant may fail an attention check, a sensor may malfunction, a response time may be physically implausible, or eligibility criteria may have been violated.

Now consider a different process. The full sample gives p =.08. Removing participants who completed unusually quickly gives p =.06. Removing another category of “outliers” produces p =.04, and only this final exclusion rule appears in the paper.

The exclusion itself might even sound defensible. What matters is whether the rule was selected because of its effect on the result. When several plausible inclusion rules are available, outcome-dependent selection among them can create a questionable use of analytical flexibility.

Searching Across Subgroups Until an Effect Appears

A treatment may show no clear effect in the full sample. Researchers then examine men and women separately, younger and older participants, high- and low-performing participants, several regions, different diagnostic groups, or combinations of these characteristics.

Eventually, one subgroup produces p <.05.

A scientifically motivated subgroup analysis can be perfectly legitimate. Searching many subgroups and presenting only the successful one as compelling evidence is different. As the number of investigated possibilities grows, so does the opportunity to encounter an apparently interesting pattern by chance.

The distinction becomes especially important when deciding when subgroup analysis turns into cherry-picking.

Changing How a Variable Is Defined

Many constructs can be operationalized in more than one defensible way. Achievement might be represented by a final examination score, course grade, standardized score, or composite measure. Age could be continuous or categorized. A scale might use all items or exclude an item with poor psychometric performance.

Alternative operationalizations can be useful for assessing robustness. They become problematic when researchers try several definitions, discover which produces the preferred inferential result, and report that definition as though it were the obvious choice all along.

Trying Different Statistical Models Until One Crosses the Threshold

The same research question can sometimes be analyzed using several plausible model specifications. Researchers might test different interaction terms, transformations, random-effects structures, estimators, or sets of predictors.

Again, model comparison is not itself p-hacking. Statistical modeling often requires examining alternatives. The danger lies in treating p <.05 as the model-selection criterion and then presenting the winning specification without revealing the search.

This is why trying several analyses until one works requires more nuance than simply counting how many models were run.

Collecting More Data After Checking the P-Value

Suppose a researcher initially collects 100 observations and obtains p =.08. Rather than following a predetermined stopping rule, the researcher adds 20 participants and checks again. The result becomes p =.06. Another 20 are collected, producing p =.04, and data collection stops.

This practice is commonly called optional stopping. Under conventional fixed-sample testing, repeatedly looking at the accumulating data and allowing the observed result to determine whether collection continues can alter the error properties of the procedure.

This does not mean researchers can never analyze accumulating data. Sequential and adaptive designs can formally permit interim analyses when the stopping procedure and statistical method appropriately account for them. The problem is unplanned peeking coupled with a significance-dependent decision to stop.

The Researcher May Not Feel Like They Are “Hacking” Anything

The term p-hacking can imply deliberate manipulation, but the underlying statistical problem does not require conscious dishonesty. Gelman and Loken described a related “garden of forking paths” problem in which analytical decisions are influenced by features of the observed data even when the researcher does not literally run every possible analysis.

A researcher might sincerely believe that an exclusion became appropriate after inspecting the data, that a subgroup suddenly became theoretically interesting, or that one operationalization now seems more sensible. Those judgments may even be reasonable. Yet if the observed data influence which inferential path is taken, the conventional interpretation of the final test may no longer correspond neatly to a single analysis selected independently of the data.

Watch Out

Intent is not a statistical correction. A researcher can introduce result-dependent flexibility without deliberately trying to deceive anyone. The relevant question is how the analytical procedure behaved, not merely whether the researcher felt they were “fishing.”

04 · A Practical Example

How a Nonsignificant Result Can Gradually Become “Significant”

Hypothetical Example

A Digital Learning Intervention

A researcher tests whether a digital learning intervention improves examination performance. The original plan is to compare intervention and control groups after 120 students have completed the study.

Planned analysis: p =.11 The primary comparison does not reach the conventional.05 threshold.
Add a covariate: p =.08 The researcher adjusts for students' previous grades.
Exclude several low-engagement participants: p =.06 A plausible engagement criterion is introduced after inspecting the data.
Analyze only first-year students: p =.03 The effect becomes statistically significant in one subgroup.
Report the successful analysis The manuscript highlights the first-year subgroup and does not make the preceding analytical search clear.

No single step necessarily proves p-hacking. There could be substantive reasons for covariate adjustment, exclusions, and subgroup analysis. The problem becomes apparent from the selection mechanism: each new choice is evaluated partly by whether it moves the result toward statistical significance, and the eventual successful result is detached from the unsuccessful paths that led to it.

A more transparent analysis could report the planned result first, identify subsequent analyses as exploratory or sensitivity analyses, explain why they were performed, and show how conclusions change across reasonable specifications. The reader can then see that the apparent effect depends on particular analytical decisions.

05 · What Researchers Often Get Wrong

What P-Hacking Is Not

Misconception

Running More Than One Analysis Means You P-Hacked

No. Researchers routinely compare models, perform sensitivity analyses, diagnose assumptions, examine robustness, and conduct exploratory analyses. Multiple analyses become concerning when favorable results determine what is ultimately treated as the evidential result without appropriate disclosure or statistical treatment.

Misconception

Changing an Analysis After Seeing the Data Is Always Wrong

Data can reveal violated assumptions, measurement problems, coding errors, influential observations, or other issues that genuinely require changes. The appropriate response is not to preserve an unsuitable analysis for the sake of purity. Explain the change, justify it, and distinguish it from the original plan when that distinction matters.

Misconception

Only Dishonest Researchers P-Hack

Outcome-driven analytical choices can arise through motivated reasoning, ambiguity about the appropriate analysis, or ordinary responses to patterns in the data. The National Academies notes that researchers may p-hack without recognizing the consequences. Statistical distortion does not require a villain twirling a methodological moustache.

Misconception

A Justifiable Final Analysis Erases the Search That Produced It

A model can be defensible in isolation while its inferential interpretation still depends on how it was selected. If many plausible specifications were considered and the reported one was chosen because it yielded the desired result, readers need enough information to understand that selection process.

Misconception

P-Hacking Only Matters When P Is Barely Below.05

The underlying issue is data-dependent analysis and selective inference, not a particular final decimal. A result below.01 is not automatically immune if it emerged from a sufficiently broad or adaptive search. Conversely, a p-value just below.05 does not itself prove that p-hacking occurred.

06 · What This Means for You

How to Tell Whether Your Own Analysis Is Drifting Toward P-Hacking

When you change an analysis, ask what prompted the change. Would you have made the same decision if the previous result had already supported your hypothesis?

That counterfactual question is surprisingly useful. If you would exclude an observation only when including it produces p =.06, but happily retain it when the result is p =.03, statistical significance is influencing a methodological decision that should ordinarily have another justification.

A simple decision framework

If the alternative analysis tests robustness
Show whether the substantive conclusion changes across reasonable specifications rather than reporting only the preferred version.
If the data reveal a genuine methodological problem
Correct the analysis, document the reason, and disclose consequential departures from the original plan.
If you are exploring an unexpected pattern
Explore it, but identify the analysis as data-informed or exploratory rather than presenting it as a prespecified confirmatory test.
If you keep changing decisions because the result is nonsignificant
Stop treating statistical significance as the criterion for analytical choice. Reconsider the analysis on methodological grounds and disclose the search already conducted.

Preregistration or another prospective analysis plan can reduce ambiguity by documenting important decisions before the relevant results are known. It does not prohibit later exploration. The Center for Open Science explicitly distinguishes planned analyses from additional exploratory work and encourages transparent reporting of both.

07 · A Quick Checklist

Check Whether Your Analysis Shows Signs of P-Hacking

Before treating a result as confirmatory, check:
Did statistical significance influence which outcome you chose to emphasize?
Did you try alternative covariates, exclusions, transformations, or models and retain the version producing the preferred result?
Did you search multiple subgroups and report only the subgroup showing an effect?
Did you repeatedly check significance while collecting data and allow the result to determine when you stopped?
Would you have made the same analytical decisions if the earlier analysis had already supported your hypothesis?
Have you disclosed consequential analytical alternatives and deviations from your original plan?
Have you distinguished exploratory analyses from genuinely prespecified confirmatory analyses?
Does your interpretation reflect the full analytical search rather than only the most favorable result?
08 · Frequently Asked Questions

Frequently Asked Questions About P-Hacking

Is removing an outlier p-hacking?

Not by itself. Outlier exclusions can be methodologically justified. Concern arises when researchers try different exclusion rules, inspect how each affects the result, and select the rule because it produces a favorable inferential outcome without transparently reporting that process.

Is adding a control variable p-hacking?

No. Covariate adjustment may be theoretically or statistically appropriate. It becomes problematic when researchers search across covariate combinations primarily to obtain a preferred result and selectively report the successful specification.

Is analyzing subgroups p-hacking?

Not inherently. Prespecified or scientifically justified subgroup analyses can be informative. Searching numerous subgroups until one produces statistical significance and presenting only that result creates a different inferential problem.

Can p-hacking happen without running dozens of tests?

Yes. Data-dependent decisions can create what Gelman and Loken describe as a garden of forking paths even when researchers do not consciously calculate every possible alternative. The analysis ultimately performed may depend on patterns observed in the data.

Does preregistration completely prevent p-hacking?

No. Preregistration can document planned decisions and make later deviations more visible, but researchers still need appropriate methods, complete reporting, and sound interpretation. It is a transparency mechanism, not a guarantee that an analysis is correct.

What should I do if I already tried many analyses?

Do not pretend the successful analysis was the only one considered. Document the consequential analytical search, distinguish exploratory findings from prespecified tests, examine robustness across reasonable specifications where appropriate, and temper inferential claims accordingly.

09 · The Bottom Line

P-Hacking Is Usually About the Process, Not One Statistical Test

The Bottom Line

P-hacking occurs when flexibility in collecting, analyzing, or selecting data is used in a result-dependent search for favorable statistical evidence and that search is not appropriately reflected in the analysis or reporting.

Changing a model, removing an observation, testing a subgroup, adding a covariate, or conducting another analysis is not automatically improper. Look at the sequence of decisions: what motivated each choice, how many plausible paths existed, what determined which result was emphasized, and whether readers are shown enough of that process to interpret the evidence accurately.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes