Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Is Trying Several Analyses Until One Works Always P-Hacking?

Trying several analyses is not automatically p-hacking. The crucial distinction is whether alternatives answer legitimate analytical questions or are searched and selected because one produces the desired result.

486
Are Multiple Analyses P-Hacking? Guide 486 of 530
01 · The Question

If You Try More Than One Analysis, Have You P-Hacked?

Your first analysis does not always have to be your last. Perhaps a model assumption fails. A reviewer requests another specification. You want to know whether an effect survives a different exclusion rule. Or you deliberately compare several plausible models because no single specification is unquestionably correct.

Then one analysis produces a statistically significant result while another does not. Have you crossed into p-hacking simply because you tried more than one?

No. Counting analyses is the wrong diagnostic. What matters is why the analyses were performed, how they were selected, what inferential claims are made from them, and what readers are ultimately shown.

02 · The Short Answer

Multiple Analyses Are Not Automatically P-Hacking

In Brief

No. Trying several analyses is not automatically p-hacking. Multiple analyses can be appropriate for exploration, assumption checking, model comparison, sensitivity analysis, robustness assessment, or answering distinct research questions.

The practice becomes problematic when researchers use observed results to search across analytical choices for a preferred finding, then selectively present that finding or interpret it as though it came from a single prespecified test without accounting for the search.

03 · What You Need to Know

The Difference Between Analytical Investigation and Significance Hunting

Researchers Often Have More Than One Defensible Analysis

Statistical analysis is rarely a matter of pressing one uniquely correct button. Researchers may have legitimate choices about variable coding, functional form, missing-data procedures, covariates, exclusion criteria, estimators, distributional assumptions, or model structure.

Sometimes those alternatives are part of the scientific question. You might want to know whether a finding is robust to several defensible specifications. In other cases, diagnostic information reveals that the original model is inappropriate and should be revised.

The existence of researcher degrees of freedom therefore does not mean researchers should pretend that uncertainty about analysis does not exist. The more useful goal is to manage that flexibility without allowing favorable outcomes to determine which analytical path receives privileged status.

Purpose Matters: Are You Testing Robustness or Searching for Success?

Consider two researchers who each run six regression models.

Researcher A Researcher B
Defines six theoretically defensible specifications to assess robustness Keeps modifying the model after each nonsignificant result
Examines whether conclusions are consistent across them Stops when one specification gives p <.05
Reports the consequential specifications and their differences Reports only the successful specification
Interprets instability as uncertainty about the finding Treats the successful result as decisive evidence

Both researchers ran multiple analyses. Their inferential procedures are not equivalent.

For Researcher A, variation across analyses is information. For Researcher B, variation is an obstacle to be searched until a desirable answer appears.

Robustness Analysis Asks Whether the Conclusion Survives Reasonable Choices

A sensitivity or robustness analysis intentionally examines what happens when reasonable analytical decisions change. For example, a researcher might estimate a relationship using alternative defensible measures, compare results with and without influential observations, or assess several plausible model specifications.

If the substantive conclusion remains similar, that stability may strengthen confidence that it is not an artifact of one arbitrary choice. If the conclusion changes substantially, that instability is itself scientifically important.

P-hacking reverses this logic. Rather than asking, “Does my conclusion survive reasonable alternatives?”, the implicit question becomes, “Which alternative gives me the conclusion I want?”

Robustness checking Uses alternative reasonable analyses to learn whether conclusions depend on analytical choices.
Result-driven specification searching Uses alternative analyses to locate a preferred result and privileges that result because it is favorable.

Exploratory Analysis Is Also Legitimate

Researchers are allowed to explore. Indeed, exploratory analysis can reveal unexpected patterns, generate hypotheses, identify nonlinear relationships, uncover heterogeneity, and expose problems with measurements or models.

The Center for Open Science explicitly recognizes exploratory and confirmatory analyses as valuable while emphasizing the importance of distinguishing between them. If you explore 30 relationships and discover an interesting association, that finding can be reported. What changes is the strength and type of inference you can reasonably make from the same data that generated the discovery.

You should therefore report exploratory findings as exploratory rather than reconstructing the analysis as though the successful relationship had been uniquely predicted beforehand.

Data-Dependent Decisions Can Matter Even If You Run Only One Final Test

There is a subtler complication. P-hacking and related problems cannot always be identified by asking how many analyses appear in your software history.

Gelman and Loken's “garden of forking paths” illustrates how an analysis can depend on observed data even when a researcher does not consciously run dozens of tests. Suppose you notice that an effect appears especially pronounced among older participants and, because that pattern looks theoretically meaningful, decide that age should define the focal subgroup. Had a different pattern appeared, you might have chosen a different analysis.

Only one focal test may ultimately be computed, but the path to that test was informed by the data. The set of potential analyses therefore matters, not merely the number of commands actually executed.

Watch Out

“I only ran one final test” does not necessarily establish that the test was selected independently of the data. Conversely, “I ran ten models” does not establish p-hacking. The analytical decision process matters more than the raw count.

Multiple Testing and P-Hacking Are Related but Not Identical

If you conduct many statistical tests, the probability of obtaining at least one apparently significant result by chance can increase, depending on the tests and inferential framework. This is a multiplicity problem even when every test is reported honestly.

P-hacking adds a selection problem: analytical choices are searched or adapted based on results, and favorable outcomes receive preferential treatment.

Researchers have several ways to address multiplicity depending on the design and purpose of the analysis. These may include limiting prespecified confirmatory tests, using appropriate multiplicity adjustments, modeling related effects jointly, emphasizing estimation and uncertainty rather than binary significance decisions, and clearly separating exploratory analyses from confirmatory claims.

The appropriate solution depends on the statistical problem. There is no universal instruction to “correct every p-value whenever more than one analysis exists.”

Changing an Analysis Can Be Methodologically Necessary

Suppose your prespecified linear model shows severe assumption problems. Keeping it simply because it was planned would not rescue the inference. You may need another model.

Similarly, you might discover a coding error, an invalid measurement, an unforeseen pattern of missingness, or an observation that should have been excluded under the study's substantive eligibility rules.

The appropriate response is to make the defensible change and document it. The National Academies recommends describing analytical decisions, decisions about data inclusion and exclusion, when consequential decisions were made, and whether analyses were exploratory or confirmatory.

Transparency should not become a methodological straitjacket. The goal is not obedience to an analysis plan after it becomes demonstrably unsuitable. It is an accurate account of how the final inference was reached.

Selective Reporting Is Often What Makes the Search Especially Misleading

Suppose you estimate eight reasonable models and they produce materially different results. Reporting only the strongest model deprives readers of information about the finding's dependence on specification.

This is why p-hacking frequently overlaps with selective reporting and other questionable research practices. The analytical search creates opportunities for a favorable result, while selective reporting makes that result look less contingent than it really is.

The American Statistical Association's Task Force on Statistical Significance and Replicability has similarly cautioned that model choice, insufficient description of analytical procedures, and selection of results to report can contribute to replicability problems. It also notes that even highlighting a few persuasive findings among a larger set may distort the apparent evidence.

04 · A Practical Example

The Same Six Analyses Can Tell Two Very Different Methodological Stories

Hypothetical Example

Does Online Discussion Improve Academic Performance?

A researcher estimates the association between participation in online discussions and final examination scores. Several defensible decisions exist about prior-achievement adjustment, attendance, and treatment of highly influential observations.

Scenario A: Robustness Checking

Define plausible specifications The researcher identifies several substantively defensible models and explains what each tests.
Compare the results The estimated association is positive in all six models, although uncertainty and effect estimates vary.
Report the variation The researcher presents the principal analysis alongside consequential alternatives and explains the degree of robustness.
Interpret accordingly The conclusion reflects the full pattern rather than whichever model has the smallest p-value.

Scenario B: Significance Hunting

Model 1: p =.14 The researcher adds another covariate.
Model 2: p =.09 A different exclusion rule is tried.
Models 3–5: still nonsignificant Additional combinations are tested.
Model 6: p =.047 The search stops.
Only Model 6 is emphasized The manuscript presents this specification as the analysis supporting the claimed association.

The number of analyses is identical. What differs is the role the observed results played in selecting and interpreting them.

In the first scenario, multiple analyses reveal how much the conclusion depends on modeling choices. In the second, multiple analyses function as repeated opportunities to obtain a preferred threshold-crossing result.

05 · What Researchers Often Get Wrong

Common Mistakes When Judging Multiple Analyses

Misconception

The First Analysis Is Automatically the Most Valid

No. Your first model can be misspecified, based on an error, or poorly suited to the observed data structure. Methodological validity depends on the design, assumptions, research question, and analytical reasoning, not chronological priority alone.

Misconception

Once You Preregister, You Cannot Change Anything

You can deviate from a preregistered analysis when there is a sound reason. The important step is to disclose consequential deviations and distinguish them from the original plan. Additional exploratory analyses are also compatible with preregistration when clearly identified as such.

Misconception

If Every Model Is Defensible, You Can Report Whichever One You Prefer

Individual plausibility does not remove selection effects. If several defensible specifications produce different conclusions, choosing one because its result is favorable conceals meaningful analytical uncertainty. The disagreement among specifications may itself be part of the finding.

Misconception

You Have to Put Every Experimental Analysis in the Main Paper

No. Research reporting requires judgment, and exploratory work can generate extensive output. The aim is to disclose the analyses and decisions necessary for readers to understand how the evidential claim was produced. Supplementary materials, repositories, analysis code, and robustness summaries may help when full detail would overwhelm the main text.

Misconception

If You Did Not Intend to Find Significance, There Is No Statistical Problem

Intent and statistical operating characteristics are different questions. Data-dependent analytical decisions can affect inference even without a conscious search for p <.05. This is one reason transparent reporting of the decision process matters.

06 · What This Means for You

What to Do When Several Analyses Are Reasonable

Do not choose between “run only one analysis” and “run everything and keep the nicest result.” Neither is a satisfactory general strategy.

Instead, identify what each analysis is doing and what question the variation among analyses can answer.

A simple decision framework

If one analysis was clearly prespecified as the primary confirmatory test
Report that result regardless of whether it crosses a significance threshold, alongside justified deviations or supplementary analyses.
If several specifications are genuinely plausible
Examine whether the substantive conclusion is robust across them rather than choosing solely by statistical significance.
If diagnostics show the original analysis is inappropriate
Use a more appropriate analysis and explain why the change was necessary.
If an alternative was suggested by an interesting feature of the observed data
Treat the resulting inference as data-informed or exploratory where appropriate and report how the analysis arose.
If results conflict across reasonable analyses
Treat that instability as substantive information rather than hiding the specifications that weaken the preferred conclusion.

A useful warning sign is the phrase, “Let's try one more thing” repeated only after unfavorable results. Sometimes one more analysis really is necessary. But if the search reliably ends when p drops below.05, the stopping rule is telling you more about the analytical process than the methodological rationale is.

07 · A Quick Checklist

Before Choosing Among Multiple Analyses

When several analyses are possible, check:
What substantive or statistical reason justifies each alternative analysis?
Were any primary analyses or decision rules specified before you examined the relevant results?
Are you using alternative analyses to assess robustness or mainly to search for statistical significance?
Would you still choose the same specification if it produced a less favorable result?
Do reasonable specifications lead to materially different conclusions?
Have you considered multiplicity when several inferential tests address the same claim?
Are consequential deviations and data-dependent decisions clearly disclosed?
Does your manuscript distinguish prespecified confirmatory analyses from later exploratory analyses?
08 · Frequently Asked Questions

Frequently Asked Questions About Running Multiple Analyses

How many analyses can I run before it becomes p-hacking?

There is no numerical cutoff. Running two analyses can be problematic if the preferred one is selected because of its result, while dozens of transparently reported simulations, sensitivity analyses, or model checks may be methodologically appropriate. Purpose, selection, inference, and reporting matter more than the count.

Can I try another analysis if my first result is nonsignificant?

Yes, if there is a legitimate reason. A nonsignificant result by itself is not a methodological reason to keep changing the analysis until something becomes significant. Explain why the alternative is appropriate and whether it was motivated by the observed data.

What if the alternative analysis is statistically better than my original one?

Use the method that is defensible for the data and research question. If the change occurred after inspecting the results, disclose that fact when it affects interpretation and explain why the revised method is preferable rather than pretending it was always the planned analysis.

Should I report every model I tried?

Not necessarily every trivial experiment, but readers should receive enough information to understand consequential analytical flexibility and assess whether the reported conclusion depends on selective specification. Material alternatives can often be summarized in sensitivity analyses or supplementary materials.

What if all reasonable analyses give approximately the same answer?

That consistency can be informative. Report the robustness assessment appropriately rather than selecting one analysis because it has the smallest p-value. Similar substantive conclusions across defensible specifications may be more meaningful than which model crosses an arbitrary threshold most comfortably.

What if only one reasonable specification is statistically significant?

Do not automatically privilege it. Investigate why the results differ across specifications and report that dependence. A conclusion that changes when a reasonable analytical choice changes should generally be presented as less robust than one that remains stable.

Does labeling analyses exploratory solve the problem?

It addresses an important transparency issue but does not make every analysis statistically appropriate. Exploratory work still requires defensible methods and cautious interpretation. The label tells readers that the analysis was not an independent prespecified test of a hypothesis generated before examining the relevant data.

09 · The Bottom Line

Several Analyses Can Strengthen Research Rather Than Undermine It

The Bottom Line

Trying several analyses is not inherently p-hacking; the problem arises when researchers search among analytical choices for a preferred result and then report or interpret that result without appropriately reflecting the search that produced it.

Alternative analyses can reveal robustness, expose model dependence, diagnose problems, and generate new hypotheses. Treat variation among reasonable analyses as information. If your conclusion survives it, say so. If your conclusion depends heavily on one particular choice, that dependence belongs in the research story too.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes