Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

P-Hacking, HARKing, and Selective Reporting: Research Practices You Need to Avoid

P-hacking, HARKing, and selective reporting can make research evidence appear stronger or more confirmatory than it really is. Learn how these practices differ, where legitimate exploration ends, and how transparent reporting protects the integrity of your findings.

484
P-Hacking, HARKing, and Selective Reporting Guide 484 of 530
01 · The Question

When Does Flexible Research Become Misleading Research?

Real research rarely unfolds exactly as planned. You may discover an unexpected pattern, realize that another statistical model is more appropriate, examine a subgroup you had not considered, or develop a new hypothesis after seeing the data. None of those actions is automatically improper.

The problem begins when the published account conceals how those decisions were made. If you try analyses until a favorable result appears, present a hypothesis developed after seeing the results as though you predicted it beforehand, or report only the outcomes that support the preferred story, readers receive a distorted picture of the evidence.

Three terms appear frequently in discussions of these problems: p-hacking, HARKing, and selective reporting. They overlap, but they are not interchangeable. Understanding the difference matters because the remedy is not to eliminate exploration from research. It is to prevent exploration and analytical flexibility from masquerading as stronger confirmatory evidence than they actually provide.

02 · The Short Answer

What Are P-Hacking, HARKing, and Selective Reporting?

In Brief

P-hacking involves exploiting analytical or data-collection flexibility in pursuit of statistically significant results; HARKing involves presenting hypotheses developed after seeing the results as though they were specified beforehand; and selective reporting involves presenting only a subset of outcomes, analyses, or findings in a way that gives an incomplete or misleading account of the evidence.

These practices are related because each can make evidence look more convincing than the underlying research process warrants. Exploring data, testing alternative analyses, or developing new hypotheses after observing results can be legitimate, however, when those activities are methodologically justified and reported transparently.

03 · What You Need to Know

How These Practices Distort the Research Record

P-Hacking Exploits Analytical Flexibility in Search of a Significant Result

Researchers routinely make analytical decisions. They may need to decide which observations to exclude, whether to include covariates, how variables should be operationalized, which statistical model is appropriate, whether transformations are needed, and which subgroups warrant investigation. Such choices are often called researcher degrees of freedom.

Having these choices is not itself a problem. The problem arises when choices become contingent on obtaining a desirable result and that contingency is hidden from the reader.

For example, suppose an analysis produces p =.08. A researcher then removes several observations and obtains p =.06, adds a covariate and obtains p =.04, and reports only the final model. The reported analysis may look like a single planned test even though it was effectively selected from several alternatives because it crossed the significance threshold.

This is why what p-hacking looks like in practice can be subtler than simply manipulating data. Simmons, Nelson, and Simonsohn demonstrated how undisclosed flexibility involving dependent variables, sample size, covariates, and subsets of conditions can substantially increase the probability of obtaining false-positive findings.

Importantly, trying several analyses is not automatically p-hacking. Sensitivity analyses, robustness checks, model comparisons, exploratory analyses, and analyses required to investigate violated assumptions may all be legitimate. What matters is why the alternatives were examined, how inferential uncertainty is handled, and whether the analytical process is reported honestly.

Watch Out

A nominal p <.05 does not preserve a 5% false-positive rate when a researcher repeatedly tries different tests, outcomes, specifications, stopping points, or subsets and reports whichever analysis succeeds as though it were the only test performed.

HARKing Rewrites When a Hypothesis Was Developed

HARKing stands for Hypothesizing After the Results are Known, a term introduced by Norbert Kerr. The critical feature is not simply developing a hypothesis after examining data. Scientific discoveries often generate new hypotheses. HARKing occurs when a post hoc hypothesis is presented in a research report as though it had been an a priori hypothesis.

Imagine that you predicted that intervention A would improve outcome X. The result for X is null, but you unexpectedly find an association with outcome Y. You develop a plausible theoretical explanation for Y and write the manuscript as though the study had been designed from the beginning to test that prediction. The final paper erases the distinction between prediction and discovery.

That distinction matters because evidence used to generate a hypothesis is not equivalent to an independent test of that same hypothesis. A pattern discovered in one dataset may reflect a genuine phenomenon, but it may also reflect sampling variability or the many possible patterns that could have attracted attention. A new dataset can provide a stronger confirmatory test.

The appropriate response is therefore not to suppress unexpected findings. Researchers can report exploratory findings without pretending they were predicted. The problem is the retrospective reconstruction of the research story.

Selective Reporting Changes Which Evidence Readers Get to See

Selective reporting occurs when the reported evidence is chosen in a way that produces an incomplete or distorted representation of what was investigated or found. The selection can occur at several levels.

A study might measure five outcomes but emphasize only the one that produces a favorable result. A researcher might conduct several statistical models but report only the successful specification. A manuscript might describe a subgroup effect while leaving unreported the unsuccessful subgroup analyses that led to it. In intervention research, prespecified outcomes may be omitted, altered, or replaced in the published report.

Not every omission is selective reporting in the problematic sense. Papers have finite space, some analyses may genuinely be redundant, and exploratory work can generate more output than belongs in a journal article. The central issue is whether what has been omitted would materially change a reasonable reader's interpretation of the evidence.

This is particularly important when deciding whether to leave a nonsignificant analysis out of a paper. Statistical significance should not itself determine whether a relevant analysis deserves to be reported.

The Three Practices Can Occur Together

Practice What changes? Core problem Illustrative pattern
P-hacking Data collection or analytical decisions Flexibility is used to obtain a desired statistical result without adequately accounting for or disclosing that search Trying several specifications and reporting the one producing p <.05
HARKing The reported timing or status of a hypothesis A hypothesis informed by the observed results is portrayed as an a priori prediction Finding an unexpected association and writing it as the study's original prediction
Selective reporting The evidence presented to readers Relevant outcomes, analyses, or results are omitted or disproportionately presented so that the report misrepresents the evidence Reporting the significant outcome while concealing relevant null outcomes

These practices can form a sequence. A researcher might try numerous analyses until one produces a statistically significant result, construct a hypothesis that explains that result, and then publish only the successful analysis and newly constructed hypothesis. P-hacking affected the analytical search, HARKing affected how the hypothesis was portrayed, and selective reporting affected what the reader ultimately saw.

The Central Problem Is Hidden Data-Dependent Decision-Making

Many analytical choices must legitimately be made after researchers learn something about their data. A distribution may violate model assumptions. Measurement problems may become apparent. Missing data may require an alternative strategy. An unexpected result may suggest a scientifically valuable follow-up analysis.

Accordingly, a simple rule such as “never change your analysis after seeing the data” is not realistic. A more useful question is whether a decision was data-dependent, whether it affects the strength or interpretation of the inference, and whether readers are told enough to understand that dependency.

This is where analytical flexibility can become a questionable research practice. Flexibility becomes especially concerning when researchers repeatedly exploit it in the direction of preferred findings while presenting the eventual choice as inevitable or prespecified.

Exploratory and Confirmatory Research Serve Different Purposes

Exploratory analysis Uses data to discover patterns, generate hypotheses, investigate possibilities, or identify questions for further study.
Confirmatory analysis Evaluates hypotheses or analytical expectations specified independently of the particular results being used as the test.

Both are valuable. Problems arise when the distinction disappears in reporting.

An exploratory finding may be surprising, theoretically interesting, and worthy of publication. Labeling it exploratory does not make it scientifically inferior. It tells readers how the finding arose and therefore how strongly it should be interpreted as a test of a prior prediction.

Conversely, merely calling an analysis “confirmatory” does not make it rigorous. Its credibility also depends on study design, measurement, statistical assumptions, data quality, adherence to the analysis plan, and appropriate interpretation.

Preregistration Can Help, but It Does Not Replace Judgment

One way to make the distinction between planned and data-dependent decisions more visible is to document hypotheses, outcomes, exclusion criteria, sampling decisions, and analysis plans before observing the relevant results. Depending on the field and study design, this may take the form of a protocol, preregistration, statistical analysis plan, or registered report.

Prespecification creates a record against which later deviations can be identified. It does not mean researchers are forbidden from changing course. Unexpected complications are normal. A transparent report can explain what changed, why it changed, and which analyses were prespecified versus added later.

Nor does preregistration guarantee good research. A poorly designed prespecified analysis remains poorly designed. Its main value here is transparency about the temporal order of decisions, not a magical conversion of methodological choices into correct ones.

04 · A Practical Example

How an Ordinary Analysis Can Turn Into a Misleading Research Story

Hypothetical Example

A Study of a New Teaching Strategy

Suppose a researcher compares a new teaching strategy with conventional instruction. The study measures examination scores, course engagement, academic confidence, attendance, and satisfaction. The researcher initially expects the intervention to improve examination scores.

Original result The planned comparison of examination scores is not statistically significant.
Additional analyses The researcher tests the other outcomes, several covariate combinations, and multiple student subgroups. One analysis shows a statistically significant improvement in academic confidence.
New explanation After seeing this result, the researcher develops a plausible theory explaining why the teaching strategy should improve academic confidence.
Selective manuscript The paper emphasizes the confidence result, omits most unsuccessful analyses, and introduces the confidence hypothesis as though it had motivated the study from the beginning.

Several different integrity problems now coexist. Searching through analytical alternatives for a favorable significance result may constitute p-hacking, depending on how the analyses were conducted and interpreted. Presenting the newly developed confidence hypothesis as a prior prediction is HARKing. Omitting relevant outcomes and unsuccessful analyses can become selective reporting if those omissions materially distort the overall evidence.

The same unexpected confidence result could have been reported responsibly. The researcher could state that the primary analysis did not support the original hypothesis, identify the confidence analysis as exploratory, disclose relevant analytical multiplicity, and present the finding as a hypothesis worth testing independently.

The data have not changed. What changes is the honesty of the inference.

05 · What Researchers Often Get Wrong

Common Misunderstandings About Questionable Research Practices

Misconception

Any Analysis You Did Not Plan in Advance Is P-Hacking

No. Unplanned analyses can be legitimate exploratory analyses, robustness checks, responses to assumption violations, or investigations prompted by unexpected findings. The concern is not simply that an analysis was unplanned. It is whether analytical flexibility is exploited to obtain a preferred conclusion and whether the resulting evidential status is represented transparently.

Misconception

You Should Ignore Interesting Patterns You Did Not Predict

That would throw away one of the productive functions of empirical research. Unexpected findings can generate valuable hypotheses. The appropriate response is to identify their exploratory or post hoc origin rather than rewriting them as predictions made before the data were examined.

Misconception

A Statistically Significant Result Is Safe as Long as the Analysis Itself Is Valid

A statistical test can be technically valid in isolation while the larger inferential process is misleading. If many outcomes, models, subgroups, or stopping points were examined and only a successful result was selected, the nominal p-value does not communicate the full search process that produced that result.

Misconception

Only Deliberate Cheating Creates These Problems

Not necessarily. Researchers may make defensible-looking decisions one at a time without recognizing their cumulative effect. The National Academies notes that researchers may p-hack without understanding the consequences. Incentives favoring novel, clean, and statistically significant findings can also encourage biased decision-making without an explicit plan to deceive.

Misconception

You Must Report Every Statistical Test You Ever Ran

Transparent reporting does not require turning a paper into an archaeological record of every keystroke. What readers need is sufficient information to evaluate the evidential claim. Prespecified outcomes, material deviations, relevant alternative analyses, multiplicity, and results whose omission would change the interpretation deserve particular attention.

Misconception

These Practices Are Automatically the Same as Fabrication or Falsification

Research-integrity frameworks distinguish among categories of problematic conduct. The National Academies has described misleading statistical analysis and incomplete reporting that fall short of falsification as detrimental research practices. Particular acts can become more serious depending on what was manipulated, omitted, and intended, so institutional definitions and applicable policies matter. It is more accurate to evaluate the specific conduct than to assume every instance belongs to the same misconduct category.

06 · What This Means for You

How to Explore Your Data Without Misrepresenting the Evidence

You do not need to eliminate analytical flexibility. You need to make consequential flexibility visible and interpret the resulting evidence accordingly.

A simple decision framework

If an analysis was specified before examining the relevant results
Report it as planned or confirmatory, while disclosing material deviations from the plan.
If an analysis was prompted by patterns observed in the data
Report it as exploratory, post hoc, or hypothesis-generating where that distinction affects interpretation.
If several reasonable analytical specifications were tried
Do not select a preferred specification solely because it produces the desired significance result. Explain consequential alternatives and consider robustness or sensitivity analyses.
If a planned outcome or analysis produces a null or unfavorable result
Do not let statistical significance alone determine whether it disappears from the report.
If an unexpected result produces an interesting new hypothesis
Keep the finding, but distinguish discovery from confirmation and consider testing the hypothesis independently.

The same principle applies when results complicate the story you hoped to tell. Researchers are expected to interpret evidence, not merely dump output into a manuscript, so some selection and organization are unavoidable. The relevant boundary is whether telling a clear research story becomes misrepresentation.

A useful test is to imagine that another competent researcher receives your complete analysis history. Would that person interpret the strength, consistency, and confirmatory status of your evidence substantially differently from a reader of the paper? If so, the manuscript may be hiding information that matters.

Transparency does not make a messy study magically rigorous. It does, however, allow readers to evaluate the study that was actually conducted rather than a cleaner study reconstructed after the fact.

07 · A Quick Checklist

Before You Finalize Your Analysis and Results

Before reporting your findings, check:
Can you distinguish which hypotheses and analyses were planned before you examined the relevant results?
Have you disclosed consequential deviations from your protocol, preregistration, or original analysis plan?
Were alternative models, exclusions, covariates, outcomes, subgroups, or stopping decisions influenced by whether they produced favorable results?
Have you accounted appropriately for multiple testing or extensive analytical searching when making inferential claims?
Are unexpected findings clearly distinguished from hypotheses that genuinely preceded the analysis?
Have relevant null, negative, or contradictory results been retained when their omission would change the reader's interpretation?
Would your conclusions remain appropriately qualified if readers knew the full set of consequential analytical decisions?
Where possible, have you preserved analysis code, decision records, protocols, or other documentation that makes the analytical process auditable?
08 · Frequently Asked Questions

Questions Researchers Ask About P-Hacking, HARKing, and Selective Reporting

Is p-hacking the same as data fabrication?

No. P-hacking generally concerns the use of data-collection or analytical flexibility to obtain favorable statistical results, whereas fabrication involves making up data or results. They are different practices, although both can seriously undermine the reliability of the research record.

Is running several statistical models automatically p-hacking?

No. Multiple models may be scientifically or statistically justified. The concern increases when researchers search across models for a favorable result, selectively report the successful specification, and conceal the analytical search. Transparent sensitivity and robustness analyses serve a very different purpose.

Can I develop a hypothesis after seeing my results?

Yes. Data frequently generate new hypotheses. The important distinction is between developing a hypothesis from the current results and presenting that hypothesis as though it existed before those results were known. Post hoc hypotheses can be handled responsibly when their origin is transparent.

Do I have to report results that contradict my hypothesis?

Relevant evidence should not disappear simply because it fails to support the preferred hypothesis. What belongs in a particular manuscript depends on the study design and reporting context, but results that do not support a hypothesis may be essential for an accurate account of what the study found.

Is stopping data collection after obtaining significance p-hacking?

It can distort conventional inference when repeated significance testing determines an otherwise unplanned stopping point. Some study designs use legitimate prespecified sequential methods with statistical procedures designed for interim analyses. The issue is therefore not stopping early in every circumstance, but optional stopping without appropriate design and statistical treatment.

Is subgroup analysis a form of cherry-picking?

Not inherently. Subgroup analyses can answer important scientific questions. Concern arises when many subgroups are searched and only favorable findings are highlighted without appropriate disclosure, multiplicity considerations, or recognition that the result may be exploratory. This is why the boundary between subgroup analysis and cherry-picking depends heavily on how the analysis was motivated and reported.

Does preregistration prevent p-hacking and HARKing?

It can make undisclosed data-dependent changes easier to identify by documenting plans before results are known, but it cannot guarantee good methods or eliminate every questionable practice. Researchers may legitimately deviate from a preregistration when circumstances require it. The important step is to disclose material deviations and distinguish planned analyses from later exploration.

What if my exploratory result is more interesting than my planned result?

You can report and discuss it. Its scientific interest does not depend on pretending it was predicted. Explain how the finding arose, interpret it with appropriate caution, and consider independent confirmation if you want to make a stronger confirmatory claim.

09 · The Bottom Line

Explore Freely, but Report the Research That Actually Happened

The Bottom Line

P-hacking, HARKing, and selective reporting are problematic because they can hide the data-dependent choices behind a finding and make the resulting evidence appear more decisive, complete, or confirmatory than it really is.

Good research does not require pretending that every hypothesis, outcome, and analytical decision was perfectly anticipated. Exploration can be scientifically productive. The stronger standard is transparency: distinguish prediction from discovery, disclose consequential analytical flexibility, report evidence that materially affects interpretation, and let readers see enough of the research process to judge the claim on its actual merits.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes