01 · The Question
Cleaning Data Changes It, So When Does Changing It Go Too Far?
Raw research data are rarely analysis-ready. Researchers correct obvious entry errors, resolve duplicate records, standardize coding, investigate impossible values, address missingness, and sometimes exclude observations that cannot validly contribute to an analysis.
All of those actions can change a dataset.
That creates an uncomfortable question: if falsification can involve changing or omitting data, how is ordinary data cleaning different? The answer does not lie simply in whether the dataset changed. It lies in why the change was made, what evidence supports it, how consistently the rule was applied, and whether the resulting research record still accurately represents the research.
03 · What You Need to Know
The Boundary Is About Scientific Justification, Not Whether the Dataset Changes
Data Cleaning Is a Normal Part of Research
Data cleaning is not suspicious simply because it alters a working dataset. The NIH National Center for Advancing Translational Sciences describes data cleaning as a process for addressing issues such as duplicate records, missing vital information, and incorrect information, with the goal of making a dataset as complete and accurate as possible before analysis.
Different disciplines and study designs require different procedures. A clinical registry, laboratory experiment, qualitative coding dataset, sensor dataset, and large administrative database will not be cleaned in exactly the same way.
The common principle is that cleaning should improve the correspondence between the dataset and the evidence it is intended to represent, not improve the attractiveness of the research conclusion.
Falsification Can Include Changing or Omitting Data
The U.S. Public Health Service definition makes the potential overlap clear. Falsification includes manipulating research materials, equipment, or processes, or changing or omitting data or results such that the research is not accurately represented in the research record.
Consequently, “I was cleaning the data” is not a complete justification for a change. A researcher must still be able to explain what problem was identified, why the chosen correction was appropriate, and how the resulting record represents the research more accurately.
Legitimate data cleaning
A change is supported by evidence or a defensible processing rule and is intended to improve data quality or implement the analysis appropriately.
Improper data manipulation
A change selectively alters, suppresses, recodes, or omits information so that the resulting record gives a misleading account of the research.
Correcting a Verified Error Is Different From Replacing an Inconvenient Value
Suppose a participant's original questionnaire clearly records an age of 25, but the dataset says 250. Correcting 250 to 25 restores the dataset to the source evidence.
Now suppose the questionnaire itself says 25, but the researcher changes the value to 35 because 25 creates an inconvenient distribution. Both actions modify a cell. Only the first has evidentiary support for the new value.
This is why provenance matters. A legitimate correction should ordinarily be traceable to evidence showing why the existing entry was wrong and why the replacement is more accurate. Researchers can correct genuine data errors without turning correction into manipulation.
Deleting Duplicates Is Not the Same as Deleting Unwanted Observations
If a technical import creates two identical records for the same research event, removing the duplicate may prevent one observation from being counted twice. That can make the analytical dataset more faithful to what actually occurred.
Deleting a genuine observation because it weakens an effect is fundamentally different. In that case, the researcher is not correcting duplication. The researcher is changing which genuine evidence contributes to the reported result.
The same software command might delete both rows. The scientific meaning of the operations is entirely different.
Impossible Values Should Be Investigated Before They Are “Fixed”
Cleaning rules often flag impossible or implausible values. A recorded human height of 25 feet is an obvious example of information that requires investigation.
But identifying a value as impossible does not tell you what the correct value should be. Perhaps 25 was meant to be 2.5 feet. Perhaps it was 25 inches. Perhaps the wrong variable was imported. Perhaps the record belongs to someone else.
If the source record establishes the correct value, it can be corrected appropriately. If the true value cannot be established, converting an obvious error into a plausible number by guesswork creates a different problem. Sometimes the scientifically honest cleaned value is missing.
An Outlier Is Not Automatically Dirty Data
Outliers often receive disproportionate attention during cleaning because they are easy to flag statistically. Yet an unusual observation can be perfectly genuine.
A value may be extreme because of data-entry error, equipment failure, or contamination. It may also reveal genuine heterogeneity, a rare case, model misspecification, or an unexpected phenomenon.
Cleaning should therefore distinguish “unusual” from “erroneous.” Automatically deleting observations because they lie beyond an arbitrary threshold can remove legitimate evidence. The more specific question of when removing apparently bad data becomes falsification depends on the reason for exclusion and what that omission does to the research record.
Consistency Is One of the Strongest Tests of a Cleaning Rule
A defensible rule should ordinarily apply to comparable observations regardless of whether they support the researcher's expectations.
Suppose a researcher decides that values outside the instrument's validated measurement range are unusable. If that rule is applied to all such values, the decision may have a coherent methodological basis.
If values beyond the range are removed only when they weaken the hypothesis but retained when they strengthen it, the purported quality-control rule is not being applied consistently.
The Timing of a Cleaning Decision Matters, but It Is Not Decisive
Prespecified cleaning rules can reduce researcher discretion because the rules exist before the researcher knows exactly how each decision affects the result. Data-management plans, protocols, codebooks, standard operating procedures, and preregistrations can therefore make consequential decisions easier to evaluate.
Not every cleaning decision can be anticipated. Researchers discover unexpected coding problems, software bugs, instrument failures, and data anomalies after collection. A post hoc correction can be entirely legitimate.
The concern increases when a rule appears only after the researcher sees an undesirable result and disappears once the desired result is obtained.
Repeatedly Trying Cleaning Rules Until the Result “Works” Is Especially Risky
Consider a researcher who tries one outlier threshold and obtains a nonsignificant result. A second threshold still fails. A third removes two more observations and produces statistical significance, so the researcher keeps the third rule without disclosing the alternatives.
The individual rule might look superficially defensible when viewed in isolation. The decision process tells a different story.
Cleaning choices should be driven by the properties of the data and the research design, not by repeated searching for the version of the dataset that produces the preferred inference.
Recoding Can Also Become Manipulative
Data cleaning is not limited to deletion. Researchers may standardize categories, correct labels, reverse-code items, reconcile inconsistent responses, or transform variables.
Again, the reason matters. Correcting “Femlae” to the intended category “Female” is straightforward when the source and coding scheme support that correction. Moving inconvenient cases between categories after inspecting their outcomes is different.
If a coding rule is reconsidered after outcomes are known, the issue becomes closely connected to whether changing a coding decision after seeing the outcome remains a defensible methodological judgment.
Cleaning Should Be Reproducible Whenever Practical
A strong cleaning workflow makes it possible to determine how the analytical dataset was derived from the original data.
For computational research, scripted cleaning can be particularly useful because the transformations can be rerun and inspected. For other forms of research, change logs, coding records, version histories, laboratory documentation, or equivalent records may serve the same function.
The goal is not bureaucratic perfection. It is traceability. If someone asks why a value changed from 250 to 25 or why 12 observations disappeared between two versions, the research record should provide an answer more informative than “I cleaned it.”
Do Not Overwrite the Only Copy of Your Raw Data
Cleaning is much easier to defend when the original record remains available. Preserving raw or source data allows researchers to verify corrections, reproduce transformations, investigate anomalies, and reverse mistakes in the cleaning process.
Where legitimate changes are made, documenting research data corrections helps preserve an audit trail rather than allowing the cleaned dataset to erase its own history.
Watch Out
A cleaned dataset should not simply be the version that produces the nicest result. If your cleaning criterion changes after each analysis until the hypothesis is supported, the problem is no longer ordinary data hygiene. Preserve the alternatives and examine whether the decision is scientifically justified independently of its effect on the result.