Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

When Does Data Cleaning Become Data Manipulation?

Data cleaning becomes problematic when changes stop correcting or consistently processing the data and instead distort what the research actually found. The boundary depends on justification, consistency, provenance, transparency, and whether decisions are driven by desired results.

453
Data Cleaning vs. Manipulation Guide 453 of 530
01 · The Question

Cleaning Data Changes It, So When Does Changing It Go Too Far?

Raw research data are rarely analysis-ready. Researchers correct obvious entry errors, resolve duplicate records, standardize coding, investigate impossible values, address missingness, and sometimes exclude observations that cannot validly contribute to an analysis.

All of those actions can change a dataset.

That creates an uncomfortable question: if falsification can involve changing or omitting data, how is ordinary data cleaning different? The answer does not lie simply in whether the dataset changed. It lies in why the change was made, what evidence supports it, how consistently the rule was applied, and whether the resulting research record still accurately represents the research.

02 · The Short Answer

Cleaning Corrects or Systematically Processes Data; Manipulation Distorts Them

In Brief

Legitimate data cleaning uses scientifically defensible, consistently applied procedures to identify and address errors, inconsistencies, duplicates, missing information, or other data-quality problems. It crosses toward improper manipulation when researchers alter, omit, recode, or selectively process data in ways that make the research record inaccurate, particularly when decisions are driven by obtaining a preferred result.

No single operation, such as deleting a row or changing a value, determines the answer by itself. The provenance of the change, its methodological justification, consistency, timing, documentation, and effect on how the research is represented all matter.

03 · What You Need to Know

The Boundary Is About Scientific Justification, Not Whether the Dataset Changes

Data Cleaning Is a Normal Part of Research

Data cleaning is not suspicious simply because it alters a working dataset. The NIH National Center for Advancing Translational Sciences describes data cleaning as a process for addressing issues such as duplicate records, missing vital information, and incorrect information, with the goal of making a dataset as complete and accurate as possible before analysis.

Different disciplines and study designs require different procedures. A clinical registry, laboratory experiment, qualitative coding dataset, sensor dataset, and large administrative database will not be cleaned in exactly the same way.

The common principle is that cleaning should improve the correspondence between the dataset and the evidence it is intended to represent, not improve the attractiveness of the research conclusion.

Falsification Can Include Changing or Omitting Data

The U.S. Public Health Service definition makes the potential overlap clear. Falsification includes manipulating research materials, equipment, or processes, or changing or omitting data or results such that the research is not accurately represented in the research record.

Consequently, “I was cleaning the data” is not a complete justification for a change. A researcher must still be able to explain what problem was identified, why the chosen correction was appropriate, and how the resulting record represents the research more accurately.

Legitimate data cleaning A change is supported by evidence or a defensible processing rule and is intended to improve data quality or implement the analysis appropriately.
Improper data manipulation A change selectively alters, suppresses, recodes, or omits information so that the resulting record gives a misleading account of the research.

Correcting a Verified Error Is Different From Replacing an Inconvenient Value

Suppose a participant's original questionnaire clearly records an age of 25, but the dataset says 250. Correcting 250 to 25 restores the dataset to the source evidence.

Now suppose the questionnaire itself says 25, but the researcher changes the value to 35 because 25 creates an inconvenient distribution. Both actions modify a cell. Only the first has evidentiary support for the new value.

This is why provenance matters. A legitimate correction should ordinarily be traceable to evidence showing why the existing entry was wrong and why the replacement is more accurate. Researchers can correct genuine data errors without turning correction into manipulation.

Deleting Duplicates Is Not the Same as Deleting Unwanted Observations

If a technical import creates two identical records for the same research event, removing the duplicate may prevent one observation from being counted twice. That can make the analytical dataset more faithful to what actually occurred.

Deleting a genuine observation because it weakens an effect is fundamentally different. In that case, the researcher is not correcting duplication. The researcher is changing which genuine evidence contributes to the reported result.

The same software command might delete both rows. The scientific meaning of the operations is entirely different.

Impossible Values Should Be Investigated Before They Are “Fixed”

Cleaning rules often flag impossible or implausible values. A recorded human height of 25 feet is an obvious example of information that requires investigation.

But identifying a value as impossible does not tell you what the correct value should be. Perhaps 25 was meant to be 2.5 feet. Perhaps it was 25 inches. Perhaps the wrong variable was imported. Perhaps the record belongs to someone else.

If the source record establishes the correct value, it can be corrected appropriately. If the true value cannot be established, converting an obvious error into a plausible number by guesswork creates a different problem. Sometimes the scientifically honest cleaned value is missing.

An Outlier Is Not Automatically Dirty Data

Outliers often receive disproportionate attention during cleaning because they are easy to flag statistically. Yet an unusual observation can be perfectly genuine.

A value may be extreme because of data-entry error, equipment failure, or contamination. It may also reveal genuine heterogeneity, a rare case, model misspecification, or an unexpected phenomenon.

Cleaning should therefore distinguish “unusual” from “erroneous.” Automatically deleting observations because they lie beyond an arbitrary threshold can remove legitimate evidence. The more specific question of when removing apparently bad data becomes falsification depends on the reason for exclusion and what that omission does to the research record.

Consistency Is One of the Strongest Tests of a Cleaning Rule

A defensible rule should ordinarily apply to comparable observations regardless of whether they support the researcher's expectations.

Suppose a researcher decides that values outside the instrument's validated measurement range are unusable. If that rule is applied to all such values, the decision may have a coherent methodological basis.

If values beyond the range are removed only when they weaken the hypothesis but retained when they strengthen it, the purported quality-control rule is not being applied consistently.

The Timing of a Cleaning Decision Matters, but It Is Not Decisive

Prespecified cleaning rules can reduce researcher discretion because the rules exist before the researcher knows exactly how each decision affects the result. Data-management plans, protocols, codebooks, standard operating procedures, and preregistrations can therefore make consequential decisions easier to evaluate.

Not every cleaning decision can be anticipated. Researchers discover unexpected coding problems, software bugs, instrument failures, and data anomalies after collection. A post hoc correction can be entirely legitimate.

The concern increases when a rule appears only after the researcher sees an undesirable result and disappears once the desired result is obtained.

Repeatedly Trying Cleaning Rules Until the Result “Works” Is Especially Risky

Consider a researcher who tries one outlier threshold and obtains a nonsignificant result. A second threshold still fails. A third removes two more observations and produces statistical significance, so the researcher keeps the third rule without disclosing the alternatives.

The individual rule might look superficially defensible when viewed in isolation. The decision process tells a different story.

Cleaning choices should be driven by the properties of the data and the research design, not by repeated searching for the version of the dataset that produces the preferred inference.

Recoding Can Also Become Manipulative

Data cleaning is not limited to deletion. Researchers may standardize categories, correct labels, reverse-code items, reconcile inconsistent responses, or transform variables.

Again, the reason matters. Correcting “Femlae” to the intended category “Female” is straightforward when the source and coding scheme support that correction. Moving inconvenient cases between categories after inspecting their outcomes is different.

If a coding rule is reconsidered after outcomes are known, the issue becomes closely connected to whether changing a coding decision after seeing the outcome remains a defensible methodological judgment.

Cleaning Should Be Reproducible Whenever Practical

A strong cleaning workflow makes it possible to determine how the analytical dataset was derived from the original data.

For computational research, scripted cleaning can be particularly useful because the transformations can be rerun and inspected. For other forms of research, change logs, coding records, version histories, laboratory documentation, or equivalent records may serve the same function.

The goal is not bureaucratic perfection. It is traceability. If someone asks why a value changed from 250 to 25 or why 12 observations disappeared between two versions, the research record should provide an answer more informative than “I cleaned it.”

Do Not Overwrite the Only Copy of Your Raw Data

Cleaning is much easier to defend when the original record remains available. Preserving raw or source data allows researchers to verify corrections, reproduce transformations, investigate anomalies, and reverse mistakes in the cleaning process.

Where legitimate changes are made, documenting research data corrections helps preserve an audit trail rather than allowing the cleaned dataset to erase its own history.

Watch Out

A cleaned dataset should not simply be the version that produces the nicest result. If your cleaning criterion changes after each analysis until the hypothesis is supported, the problem is no longer ordinary data hygiene. Preserve the alternatives and examine whether the decision is scientifically justified independently of its effect on the result.

04 · A Practical Example

One Dataset, Two Very Different Cleaning Workflows

Hypothetical Example

Cleaning Examination Scores Before Analysis

A researcher has records from 300 students. Before analysis, the dataset contains several obvious anomalies.

Identify the problems Two rows are exact duplicates created during data import. One examination score is recorded as 850 even though the assessment ranges from 0 to 100. Four students have missing scores.
Defensible cleaning The researcher removes the confirmed duplicate records, checks the original scoring record and discovers that 850 should be 85, corrects the entry, and retains the four genuinely missing scores as missing. Every change is logged.
What happens next The analysis produces a smaller effect than expected. Several valid low-scoring students in the intervention group contribute to the weak result.
Outcome-driven manipulation The researcher labels those valid scores “abnormal,” removes them, reruns the analysis, and continues changing the threshold until the intervention becomes statistically significant.
The difference The first changes are tied to identifiable data-quality problems and source evidence. The later removals are driven by their effect on the desired conclusion rather than evidence that the observations are erroneous.

The dividing line is not that one workflow changes data and the other does not. Both do. The difference is what justifies those changes and whether the resulting dataset accurately represents the research.

05 · What Researchers Often Get Wrong

Common Misunderstandings About Cleaning Research Data

Misconception

Raw Data Must Never Be Changed Under Any Circumstances

The original research record should generally be preserved, but an analytical dataset may legitimately contain documented corrections, recodings, exclusions, and transformations. The important distinction is between preserving source evidence and pretending that no processing can ever occur.

Misconception

If a Value Looks Wrong, I Can Replace It With the Most Likely Value

Recognizing that a value is erroneous does not necessarily reveal the correct replacement. Recover the value from reliable evidence where possible. If the correct value cannot be established, missingness may be more accurate than a plausible guess.

Misconception

Outlier Removal Is Just Standard Data Cleaning

Sometimes an extreme observation is erroneous; sometimes it is genuine. Statistical extremeness alone does not establish invalidity. Researchers need a defensible analytical rationale rather than treating “outlier” as a synonym for “data I may delete.”

Misconception

If Software Performs the Cleaning Automatically, the Decision Is Objective

Software executes rules chosen by people. Automated thresholds, deduplication procedures, range checks, and recoding algorithms can still be inappropriate. Researchers remain responsible for understanding what the procedure does and whether its assumptions fit the data.

Misconception

If I Document a Manipulation, It Becomes Acceptable

Documentation improves transparency but cannot transform an indefensible change into a scientifically sound one. Researchers need both an appropriate rationale and an accurate record of what was done.

06 · What This Means for You

Build a Cleaning Trail You Could Defend Without Knowing the Result

A useful test is to imagine explaining each consequential cleaning decision before seeing whether it helps your hypothesis. Could you justify the same rule based on the protocol, source evidence, measurement process, or analytical principles alone?

A simple decision framework

If source evidence shows that a recorded value is wrong and establishes the correct value
Correct it while preserving the original record and documenting the source of the correction.
If a value is impossible or suspicious but the correct value cannot be recovered
Do not invent a replacement. Determine the appropriate treatment, which may include marking it missing.
If a cleaning rule identifies a class of data-quality problems
Apply the rule consistently to comparable observations rather than selectively according to their effect on the results.
If a cleaning decision changes after you inspect the outcome
Document the change, justify it independently of the desired result, and consider sensitivity analyses or independent methodological review.
If several defensible cleaning choices produce materially different conclusions
Treat that sensitivity as scientifically relevant rather than quietly selecting the most favorable version.

Whenever practical, keep the original data immutable and create cleaned or analysis-ready versions separately. Use scripts, logs, version control, change records, or equivalent documentation appropriate to your field so that consequential transformations remain traceable.

07 · A Quick Checklist

Before Calling a Change “Data Cleaning,” Check What It Actually Does

Before finalizing a cleaned dataset, check:
Can every consequential correction, exclusion, recoding, or transformation be explained scientifically or from source evidence?
Have you preserved the original or source data rather than overwriting the only available record?
Are cleaning rules applied consistently to observations that meet the same conditions?
Have you distinguished impossible or erroneous values from observations that are merely unusual?
Can another qualified researcher determine how the analytical dataset was derived from the original data?
Did any cleaning criterion emerge only after you saw how it affected the outcome?
Where alternative defensible cleaning choices exist, have you examined whether the conclusion depends on the choice?
Does the final research record accurately represent consequential cleaning decisions and exclusions?
08 · Frequently Asked Questions

Frequently Asked Questions About Data Cleaning and Manipulation

Is changing an incorrect value data manipulation?

Changing data is not automatically improper manipulation. If reliable evidence establishes that a value was entered incorrectly and establishes the correct value, a documented correction can make the dataset more accurate rather than less accurate.

Should I delete impossible values?

Investigate them first. An impossible value indicates a problem but does not necessarily tell you what caused it or what the correct value should be. Depending on the evidence and analysis, correction, exclusion, or treatment as missing may be appropriate.

Is winsorizing data falsification?

Not inherently. Winsorization is a statistical treatment that modifies extreme values. Whether it is appropriate depends on the research question, analytical assumptions, justification, and transparent reporting. Problems arise if the procedure is used selectively or concealed in a way that inaccurately represents the analysis.

Can I remove duplicate records?

Yes, when you can establish that they are genuinely duplicate representations of the same research event rather than separate observations. Document the deduplication rule and preserve enough information to verify what was removed.

Is it misconduct to change a cleaning rule after seeing the data?

Not automatically. Unexpected data problems sometimes require revised procedures. The concern is greater when the change is driven by its effect on the desired result rather than a newly identified scientific or data-quality problem. Formal misconduct findings require additional elements under the applicable policy.

Do I need to report every minor cleaning correction in the paper?

Reporting detail depends on the study and disciplinary conventions, but the research record should preserve consequential transformations and the methods should provide enough information for readers to understand important processing decisions. Internal documentation can be more detailed than the published methods section.

Can legitimate data cleaning change my statistical result?

Yes. A scientifically justified correction or exclusion can change an estimate, confidence interval, p-value, or substantive conclusion. A change in the result does not make the cleaning improper. The question is whether the cleaning decision was justified independently of obtaining that result.

09 · The Bottom Line

Clean the Data to Represent the Research Better, Not to Make the Result Better

The Bottom Line

Data cleaning remains legitimate when corrections, exclusions, recodings, and transformations have defensible scientific or evidentiary foundations and preserve an accurate research record. It crosses toward manipulation when researchers selectively alter the data or processing rules to produce a preferred representation or result.

Preserve the source data, make consequential changes traceable, apply comparable rules consistently, and pay particular attention to decisions made after outcomes are known. A good cleaning workflow can explain not only what changed, but why.

10 · Sources and Further Reading

Authoritative Sources on Data Cleaning and Falsification

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes