03 · What You Need to Know
Analysis Plans Can Change, but the Research Record Should Not Rewrite Their History
There Are Legitimate Reasons to Change an Analysis
No statistical analysis plan can anticipate every problem. Researchers may discover that a model's assumptions are untenable, a variable was incorrectly coded, an analytical procedure is inappropriate for the observed data structure, software behaves unexpectedly, or a planned model cannot be estimated.
A new analysis may also answer an interesting question that was not anticipated originally.
None of these situations automatically implies misconduct. Science would be rather brittle if researchers were forbidden from responding intelligently to what they learn during analysis.
What Changes After Seeing the Results Is the Evidentiary Meaning of the Analysis
Suppose a researcher specifies one primary analysis before examining the outcomes. The analysis produces no convincing evidence of the predicted effect. Afterward, the researcher tries several alternative transformations, covariate combinations, exclusion thresholds, subgroup definitions, and statistical models. One produces a strong effect.
That alternative analysis may reveal something scientifically interesting. But it did not arise under the same conditions as the original confirmatory test.
The researcher already knew something about the outcomes while choosing among analytical possibilities. If the final paper presents the successful alternative as though it were the original planned test, readers lose information needed to interpret the evidence appropriately.
Exploratory Analysis Is Not Inferior Research
Exploratory analysis is an essential part of research. It can identify unexpected patterns, generate hypotheses, reveal model problems, and suggest questions for future testing.
The problem is not exploration. The problem is presenting an analysis chosen after inspecting results as though it had been specified independently of those results.
Exploratory change
The analysis is developed or modified after examining the data, and its post hoc status is represented accurately.
Misleading outcome-driven change
An analysis is selected because it produces a preferred result, while unsuccessful alternatives or the original plan are concealed in a way that misrepresents the research.
Preregistration Helps Preserve the Timeline of Decisions
Preregistration or a prospective statistical analysis plan can document hypotheses, outcomes, exclusions, models, transformations, and other analytical decisions before outcomes are known. Its practical value is partly historical: readers can distinguish decisions made independently of results from decisions made after inspecting them.
Preregistration does not prohibit deviation. Research plans sometimes need revision. The more useful principle is to preserve and explain consequential deviations rather than quietly rewriting the plan after the fact.
Research on preregistration has emphasized precisely this distinction between analyses specified before outcomes are known and those developed later. That distinction helps readers calibrate how strongly confirmatory claims are supported.
A Deviation From a Preregistration Is Not Automatically Misconduct
A preregistration is not a methodological prison sentence. Following an inappropriate analysis merely because it was preregistered can itself produce poor research.
If researchers discover a genuine problem with the planned analysis, they may have good reason to change it. A strong report can identify the original plan, explain why it became unsuitable, describe the revised analysis, and, when informative, report what the original analysis produced.
The scientific question becomes much easier to evaluate when the history is visible.
Trying Multiple Analyses Creates a Selection Problem
Many datasets support numerous plausible analyses. Researchers may choose different covariates, transformations, exclusion criteria, subgroup definitions, outcome operationalizations, or model specifications.
If analysts try enough combinations and report only the one yielding the preferred result, the reported analysis is no longer an ordinary isolated test. It has been selected from a larger analytical search.
This practice is often discussed under terms such as p-hacking or specification searching. Those concepts are broader than the formal definition of research misconduct. Not every questionable analytical practice necessarily constitutes falsification.
However, when selective analytical changes involve changing or omitting data or results such that the research is inaccurately represented, the conduct can move into the territory covered by falsification. Under the U.S. Public Health Service definition, that representation of the research record is central.
Changing Covariates After Seeing the Result Can Be Legitimate or Outcome-Driven
Suppose researchers discover that an important design variable was accidentally omitted from their model. Adding it may be methodologically justified even after seeing the initial results.
Now suppose researchers repeatedly add and remove covariates until the coefficient of interest becomes statistically significant, then report only that model without explaining how it was selected.
The operation is superficially similar: a model specification changed. The decision process is not.
Changing an Exclusion Rule Is Particularly Consequential
Analytical changes sometimes alter which observations enter the analysis. A researcher might revise an outlier threshold, redefine adherence, or change participant eligibility after seeing results.
If new evidence reveals a genuine problem, the revision may be defensible. If the threshold is adjusted because particular observations weaken the desired result, the change is much more concerning.
The boundary is closely related to whether participant exclusion remains a defensible analytical decision and whether apparently bad data are being removed for scientifically valid reasons.
Changing the Outcome Can Be More Consequential Than Changing the Statistical Test
Suppose a study identifies one primary outcome but that outcome shows little evidence of an effect. Several secondary outcomes are available, and one produces a favorable result. Reporting that secondary outcome prominently is not inherently improper.
Representing it retrospectively as though it had always been the primary outcome is different.
The issue is again historical accuracy. Readers need to know enough about what was planned and what was discovered afterward to interpret the evidentiary strength of the finding.
Reporting Both Analyses Can Often Clarify the Situation
If a planned analysis turns out to be questionable, researchers do not always need to choose between blindly following it and pretending it never existed.
Where appropriate, they can report the planned analysis and the revised analysis, explain why the latter is considered more suitable, and discuss whether the substantive conclusion depends on the analytical choice.
This kind of sensitivity analysis can turn an analytical complication into useful scientific information.
Changing an Analysis and Falsifying Research Are Not Synonyms
Under the current PHS framework, falsification involves manipulating research materials, equipment, or processes, or changing or omitting data or results such that the research is not accurately represented in the research record. Research misconduct does not include honest error or differences of opinion. A formal finding also requires a significant departure from accepted practices, the required culpable state of mind, and proof by a preponderance of the evidence.
Accordingly, a debatable modeling decision should not casually be labeled falsification. But neither should “analytical flexibility” be used as a blanket defense when the research record has been deliberately or recklessly made misleading.
Watch Out
If you have tried many analyses and know which one gives the result you want, do not write the paper as though that successful specification was the only analysis considered. The analytical search itself may be important for interpreting the result.
06 · What This Means for You
Change the Analysis When the Science Requires It, but Preserve the Analytical History
If a planned analysis becomes inappropriate, do not preserve it merely for the appearance of methodological purity. Fix the analytical problem. What you should not fix retroactively is the history of how the new analysis was chosen.
A simple decision framework
If a planned analysis remains appropriate
Use it as planned and distinguish additional analyses as appropriate.
If diagnostics or new methodological information reveal a genuine problem
Revise the analysis, document the reason, and preserve the original analytical decision and result where material.
If a new analysis emerged because you noticed an unexpected pattern
Treat it as exploratory or otherwise identify its post hoc origin rather than retrospectively presenting it as prespecified.
If you tried several specifications and selected one because it produced the preferred result
Do not conceal the specification search. Reassess the inference and consider robustness or multiverse-style analyses where appropriate.
If reasonable analytical choices lead to different substantive conclusions
Treat that instability as part of the finding rather than selecting the version that tells the most convenient story.
Changes involving classification decisions deserve similar scrutiny. If the analytical result changes because categories themselves are revised after outcomes are known, examine whether the revised coding decision remains defensible.
Whatever changes you make, preserve enough information to reconstruct the sequence. Analysis scripts, preregistrations, dated plans, version histories, decision logs, and outputs can show what was originally planned, what changed, and why.