01 · The Question
Researchers Have Choices, So When Do Those Choices Become a Problem?
Few datasets arrive with exactly one defensible analysis attached. You may need to choose how to handle missing data, whether to include covariates, how to define an outlier, which statistical model fits the data, whether a transformation is appropriate, or how to operationalize a construct.
This flexibility is not inherently a flaw. Some decisions cannot sensibly be finalized until researchers understand the characteristics of their data.
The difficulty arises when the results themselves begin steering those choices. If one decision is retained because it produces a more favorable effect, another is abandoned because significance disappears, and only the successful analytical path reaches the paper, ordinary methodological flexibility can become a questionable research practice.
03 · What You Need to Know
Where Legitimate Analytical Judgment Ends and Questionable Flexibility Begins
Researcher Degrees of Freedom Are Normal
Simmons, Nelson, and Simonsohn used the term researcher degrees of freedom to describe choices researchers make during data collection and analysis. These can include decisions about sample size, dependent variables, covariates, exclusions, and which conditions to compare.
Many other choices may arise depending on the methodology. Researchers may have to select estimators, transformations, missing-data procedures, variable definitions, model structures, time points, interaction terms, or thresholds.
Having options does not establish wrongdoing. Sometimes several approaches are genuinely reasonable. Sometimes the data reveal that an originally planned method is inappropriate. The existence of analytical judgment is therefore not the problem.
The problem is that flexibility creates multiple possible routes from the same dataset to a reported conclusion. If researchers use information about the results to favor routes producing a desired answer, conventional inferential claims can become difficult to justify.
The Motivation for a Change Matters
Suppose you planned a linear regression but diagnostic checks reveal that a key assumption is badly violated. You change the model because the original method is unsuitable. That is analytical flexibility serving methodological validity.
Now suppose the original model produces p =.08. You try a transformation and obtain p =.06. You remove several observations and obtain p =.04, so you keep that version. The analytical search is now responding to whether each choice produces the desired inferential result.
The actions may look similar from the outside. Both researchers changed an analysis. What differs is the decision criterion.
Method-driven flexibility
The analysis changes because assumptions, design considerations, measurement properties, diagnostics, theory, or other methodological evidence justify the change.
Outcome-driven flexibility
The analysis changes because an earlier result is insufficiently favorable and alternatives are evaluated partly by whether they produce the desired conclusion.
A Useful Question Is Whether You Would Make the Same Choice With a Different Result
Consider an exclusion decision. You notice that several participants have unusually high response times and consider excluding them.
Would you use the same exclusion criterion if retaining those participants gave you p =.03 instead of p =.07?
If the methodological decision changes depending on which version supports the hypothesis, the result has become part of the decision rule. That does not automatically establish intentional manipulation, but it should make you examine the analysis carefully.
This counterfactual test is useful because questionable flexibility is not always conscious. Researchers can sincerely construct plausible methodological explanations after observing which choices produce attractive findings.
The Number of Choices Matters, but There Is No Numerical Cutoff
One discretionary choice may have little consequence. Many result-dependent choices can create a large analytical search space.
Simmons and colleagues demonstrated this statistically using several ordinary researcher degrees of freedom. In their simulations, flexibility involving dependent variables, sample size, covariates, and subsets of experimental conditions substantially increased false-positive rates when researchers could exploit those choices and report favorable results.
This does not mean making four choices is questionable while making three is acceptable. There is no such threshold.
Instead, consider the space of reasonable analytical possibilities and how the final analysis was selected from that space. The more alternatives that could have been chosen based on observed results, the more cautious you should be about treating the final analysis as though it were an isolated prespecified test.
You Can Have Data-Dependent Flexibility Without Running Every Possible Analysis
Gelman and Loken's “garden of forking paths” makes the issue more subtle. Researchers do not necessarily need to run dozens of models and deliberately choose the smallest p-value.
Imagine that your data suggest a difference between younger and older participants. That observation makes age suddenly seem theoretically important, so you conduct an age-based analysis. Had the data instead suggested a difference by educational level, you might have analyzed education.
The analysis performed depends on what the dataset happened to show, even though you may have run only one follow-up test.
This means that simply asking how many analyses were actually performed does not fully capture analytical flexibility. The decision tree that could have led to different analyses also matters.
Analytical Flexibility Becomes More Concerning When It Is Hidden
A reader sees only the final paper. If that paper presents one model, one exclusion rule, and one outcome without revealing consequential alternatives, the analysis can appear much more predetermined than it actually was.
Transparency therefore changes how analytical flexibility should be interpreted. CONSORT 2025, in the context of randomized trials, recommends identifying deviations from the statistical analysis plan, explaining them, and distinguishing prespecified analyses from post hoc analyses.
This principle extends beyond clinical trials even though the specific reporting requirements differ across disciplines. Readers should be able to understand analytical decisions that materially affect the strength or meaning of the reported claim.
Robustness Checks Use Flexibility Differently
Running several analyses can actually reduce concern about arbitrary analytical decisions when the purpose is to examine robustness.
Suppose three defensible approaches to missing data produce similar effect estimates and conclusions. Showing that consistency tells readers that the finding does not depend heavily on one particular choice.
If the conclusion changes dramatically across those approaches, that is also useful information. It tells you that the evidence is analytically fragile.
Questionable flexibility does the opposite: it searches through alternatives and suppresses the instability by emphasizing whichever specification produces the preferred conclusion.
Exploration Does Not Need to Pretend to Be Confirmation
Researchers sometimes discover an analytical possibility only after inspecting their data. That can be productive rather than problematic.
You might identify an unexpected nonlinear association, a previously unconsidered subgroup, or a theoretically interesting interaction. Investigating it can generate useful knowledge.
The important distinction is evidential. If the same data suggested the analysis and provided the apparent confirmation, describe the finding accordingly. Transparent exploratory reporting allows readers to distinguish discovery from a more independent confirmatory test.
Watch Out
Preregistration is not a command to preserve a demonstrably inappropriate analysis. If circumstances justify a change, make the change. The integrity issue is whether consequential deviations are explained and whether their data-dependent nature is reflected in the interpretation.
04 · A Practical Example
Two Researchers Change the Same Analysis for Very Different Reasons
Hypothetical Example
Removing Extremely Fast Survey Responses
Two researchers studying online learning discover that several participants completed a 20-minute questionnaire in less than three minutes. Neither research team had specified a minimum completion time beforehand.
Researcher A checks data quality
The researcher examines whether extremely fast responses show evidence of careless responding, develops a defensible exclusion criterion without using statistical significance as the criterion, and reports results both with and without the affected observations.
Researcher A finds some sensitivity
The effect estimate changes slightly after exclusion. The paper explains the post hoc decision and shows readers how much the conclusion depends on it.
Researcher B checks the p-value
The original analysis produces p =.07. Several completion-time cutoffs are tried until one produces p =.04.
Researcher B reports only the successful rule
The paper describes the selected cutoff as though it were the obvious data-quality criterion and does not report the alternative analyses.
Both researchers changed their analysis after seeing the data. That fact alone does not distinguish responsible practice from questionable practice.
Researcher A used post hoc judgment to address a methodological problem and made the consequences visible. Researcher B allowed the desired statistical outcome to select the exclusion rule and then concealed the selection process.
The boundary therefore lies less in the act of changing an analysis than in the relationship among the observed result, the analytical decision, and what is subsequently reported.
06 · What This Means for You
How to Use Analytical Flexibility Without Letting the Desired Answer Choose the Method
When several analytical choices are available, separate methodological reasoning from outcome preference as much as possible.
A simple decision framework
If a choice can reasonably be made before examining the relevant results
Specify it prospectively when practical, particularly for primary confirmatory analyses.
If the data reveal a genuine methodological problem
Change the analysis when justified, document why, and report consequential departures from the original plan.
If several methods remain defensible
Use sensitivity or robustness analyses to determine whether conclusions depend on the choice.
If an analysis was suggested by the observed results
Treat its data-dependent origin as relevant to interpretation rather than retrospectively presenting it as predetermined.
If only one specification supports the preferred conclusion
Investigate and report why the specifications differ rather than automatically privileging the favorable one.
For randomized trials, current CONSORT guidance specifically asks authors to report and justify deviations from the protocol or statistical analysis plan and to clarify which analyses were prespecified and which were post hoc. It also calls for reporting how multiplicity was handled when numerous analyses are conducted.
Other research designs may use different standards, but the underlying principle travels well: your manuscript should not make an outcome-dependent analytical path look predetermined.
07 · A Quick Checklist
Check Whether Analytical Flexibility Is Becoming Questionable
When making or reporting an analytical decision, check:
What methodological reason, independent of the desired result, supports this analytical choice?
Would you probably make the same choice if the current result supported the opposite conclusion?
How many reasonable alternative specifications, exclusions, outcomes, or models were available?
Were consequential choices made before or after examining the relevant results?
Do reasonable alternative analyses materially change the conclusion?
Have important deviations from the original protocol or analysis plan been disclosed and explained?
Have data-driven exploratory analyses been distinguished from prespecified confirmatory analyses?
Would readers interpret the evidence differently if they knew the consequential analytical paths you considered?