Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Is Removing “Bad” Data Ever Falsification?

Removing data is not automatically falsification. Exclusion may be scientifically justified, but omitting observations can become falsification when it causes the research to be inaccurately represented in the research record.

452
Removing “Bad” Data Guide 452 of 530
01 · The Question

When Is It Legitimate to Remove Data From a Study?

A participant failed an attention check. An instrument malfunctioned. One measurement is physically impossible. An observation is an extreme outlier. A case does not fit the pattern the researcher expected.

Researchers sometimes describe all of these as “bad data,” but that phrase hides an important distinction. Some observations are invalid because something went wrong in measurement or data collection. Others are perfectly valid observations that simply make the results less tidy.

Removing the first kind may improve the accuracy of an analysis. Removing the second because it is inconvenient can do exactly the opposite.

02 · The Short Answer

Data Exclusion Is Defensible Only When the Scientific Reason Is Defensible

In Brief

Removing data is not automatically falsification. Researchers may legitimately exclude observations for scientifically defensible reasons, but omission can constitute falsification when data are removed so that the research is not accurately represented in the research record.

The crucial questions are why the data were excluded, whether the criterion is scientifically justified and consistently applied, when the decision was made, how sensitive the results are to the exclusion, and whether the omission is represented transparently.

03 · What You Need to Know

The Problem Is Not Removal Itself but What the Removal Does to the Research Record

Falsification Explicitly Includes Certain Omissions

Under the U.S. Public Health Service definition, falsification includes manipulating research materials, equipment, or processes, or changing or omitting data or results such that the research is not accurately represented in the research record.

The word “omitting” matters. A researcher does not need to change a numerical value to falsify a record. Selectively removing genuine observations can potentially create an inaccurate representation of the research.

But the definition does not say that every omission is falsification.

Researchers Are Sometimes Supposed to Exclude Data

Scientific analyses often require eligibility criteria, quality-control rules, handling of measurement failures, treatment of duplicate records, and decisions about observations that cannot validly contribute to a particular analysis.

The federal research misconduct policy explicitly recognized this distinction when responding to concerns about omission. It explains that appropriate omission can occur under accepted research practices, while omission becomes falsification when it misleads the reader about the research results.

That distinction prevents an absurd interpretation in which ordinary quality control becomes misconduct merely because the analyzed dataset contains fewer rows than the raw file.

“Bad Data” Is Too Vague to Justify Anything

Calling an observation “bad” tells you almost nothing scientifically.

Was the instrument outside its calibrated range? Did the participant violate an eligibility criterion? Was the observation duplicated because of an import error? Did the researcher establish that the value resulted from contamination? Or does “bad” merely mean that the observation weakens the hypothesis?

Invalid or unusable observation There is a scientifically defensible reason why the observation should not contribute to the specified analysis.
Inconvenient observation The observation is valid but undesirable because it weakens, complicates, or contradicts the preferred result.

Those categories are not always easy to determine, but researchers should not collapse them into the same label.

Outlier Status Alone Does Not Make a Data Point Invalid

An outlier is unusual relative to other observations. It is not necessarily erroneous.

A very high value might reflect a data-entry error. It might also represent genuine biological variation, an unusual participant, an unexpected social phenomenon, a rare but legitimate event, or evidence that the assumed statistical model fits poorly.

Removing an observation solely because it is extreme can therefore discard scientifically meaningful information. Researchers should investigate why the observation is unusual before deciding what analytical treatment is appropriate.

Prespecified Exclusion Rules Are Usually Easier to Defend

When exclusion criteria are established before researchers know which observations support their preferred result, there is less opportunity for the observed outcome to drive the decision.

Examples might include excluding participants who fail predefined eligibility requirements, measurements produced during documented equipment failure, or trials invalidated according to established quality-control criteria.

Prespecification does not make every rule scientifically sound. Nor does the absence of prespecification automatically make every later exclusion improper. Unexpected problems arise during research. But post hoc exclusions deserve particularly careful justification because the researcher already knows what the data look like.

A Post Hoc Exclusion Can Be Legitimate

Suppose researchers discover during analysis that a software bug duplicated several observations. The exclusion rule could not have been specified before the bug was discovered. Removing the duplicates may nevertheless make the dataset more accurate.

Similarly, an instrument malfunction may become apparent only after data collection. The relevant issue is whether there is evidence supporting the exclusion and whether the rule is applied consistently rather than opportunistically.

The broader boundary between legitimate processing and data manipulation during cleaning therefore cannot be reduced to “before analysis good, after analysis bad.”

Outcome-Driven Exclusion Is a Serious Warning Sign

Consider a researcher who runs an analysis, dislikes the result, removes one participant, reruns the model, removes another, and continues until the desired effect becomes statistically significant.

That sequence is fundamentally different from discovering that a sensor failed for a documented reason.

The problem becomes particularly serious if the exclusions are then concealed and the final dataset or report is presented as though those observations were never part of the research. Depending on the facts, such omission may cause the research to be inaccurately represented and therefore fall within falsification.

The Same Exclusion Rule Should Normally Apply to Comparable Cases

Selective enforcement can be revealing.

Suppose a researcher claims that scores more than three standard deviations from a specified reference point are invalid. If the rule removes only observations unfavorable to the hypothesis while equally extreme favorable observations remain, the stated rationale may not explain the actual decision process.

Consistency does not by itself prove that a rule is scientifically appropriate, but inconsistent application should prompt scrutiny.

Transparency Does Not Automatically Rescue a Scientifically Indefensible Exclusion

Reporting an exclusion is better than hiding it, but disclosure alone does not make the analytical choice sound.

A paper could openly state that five observations were removed because they prevented statistical significance. Transparency would allow readers to see the problem, but it would not convert the exclusion into a defensible methodological decision.

Researchers therefore need both scientific justification and accurate reporting.

Sometimes the Best Answer Is to Show the Analysis Both Ways

When a questionable observation has a plausible but uncertain reason for exclusion, a sensitivity analysis can be informative. Researchers may compare results with and without the observation and explain why one analysis is primary.

This does not solve every exclusion problem, but it prevents an ambiguous data point from silently determining the story. If the substantive conclusion changes dramatically depending on one defensible analytical decision, readers may need to know that.

Formal Falsification Requires More Than a Debatable Analytical Choice

A questionable exclusion does not automatically establish research misconduct. Under the PHS framework, misconduct excludes honest error and differences of opinion, and a finding requires a significant departure from accepted practices, the required culpable state of mind, and proof by a preponderance of the evidence.

Methodological criticism and misconduct findings are therefore not interchangeable. Researchers can make poor analytical choices without necessarily committing falsification. Conversely, describing an exclusion as an “analytical decision” does not shield deliberate or reckless manipulation that makes the research record inaccurate.

Watch Out

If your reason for removing an observation begins with “the result becomes better without it,” stop. The effect of an exclusion on your preferred conclusion is not, by itself, a scientific reason to exclude the observation.

04 · A Practical Example

The Same Data Point Can Be Removed for Very Different Reasons

Hypothetical Example

A Participant With an Extreme Score

A study includes 120 participants. Participant 73 has an unusually high outcome score that substantially weakens the experimental effect.

Scenario A: Documented measurement failure The researchers inspect the source records and discover that the instrument malfunctioned during Participant 73's measurement. The device log independently confirms the failure, and the study's quality-control procedure treats measurements produced during that failure mode as invalid.
Decision The researchers exclude the measurement, preserve the original record, document the equipment problem and exclusion, and apply the same rule to any comparable measurements.
Interpretation The exclusion has a scientific basis independent of whether the observation helps or hurts the hypothesis. Removal is not automatically falsification merely because the result changes.
Scenario B: Valid but inconvenient observation The source records show no error. The participant was eligible, the instrument worked correctly, and the observation is genuine. The researcher removes it because the treatment effect becomes statistically significant without it.
Decision The exclusion is not disclosed, and the final analysis is presented as though all valid observations were analyzed under the stated method.
Interpretation This raises a substantially different integrity problem. Omitting valid data to obtain a preferred result and thereby inaccurately representing the research can fall within the concept of falsification.
05 · What Researchers Often Get Wrong

Common Misunderstandings About Removing Research Data

Misconception

You Must Never Remove Any Observation

That is too rigid. Some observations are genuinely unusable for a particular analysis because of eligibility violations, measurement failures, duplication, contamination, or other scientifically defensible reasons. The question is whether exclusion improves the accuracy of the analysis or distorts it.

Misconception

All Statistical Outliers Should Be Deleted

Being unusual does not establish that an observation is erroneous. An outlier may contain exactly the information researchers need to understand. Investigate its provenance and scientific meaning before deciding how it should be handled.

Misconception

If Removing the Point Makes the Model Fit Better, the Removal Is Justified

Improved fit may be a consequence of removing an observation, but it does not establish that the observation was invalid. Otherwise researchers could repeatedly remove whatever the model finds inconvenient and call the resulting conformity quality control.

Misconception

Any Exclusion Decided After Seeing the Data Is Falsification

No. Researchers sometimes discover genuine data-quality problems only during analysis. A post hoc exclusion can be defensible when supported by evidence, scientifically justified, consistently applied, and transparently documented. Knowing the results does, however, create additional risk of outcome-driven reasoning.

Misconception

If I Disclose the Exclusion, It Cannot Be Falsification

Disclosure is important, but it does not automatically validate the underlying decision. A deliberately misleading analytical choice does not become scientifically defensible merely because a sentence acknowledges it.

06 · What This Means for You

Ask Why the Observation Is Invalid Before Asking Whether You Want It

The strongest exclusion decisions begin with a scientific or methodological reason, not with the direction of the result.

A simple decision framework

If an observation meets a valid prespecified exclusion criterion
Apply the criterion consistently and document the exclusion according to the study's procedures.
If a previously unknown data-quality problem is discovered
Investigate the source, formulate a scientifically defensible rule, apply it consistently, preserve the original data, and document the decision.
If an observation is merely extreme or inconvenient
Do not assume that unusual means invalid. Investigate it and consider robust methods or sensitivity analyses where appropriate.
If the exclusion was considered only after seeing that it improves the desired result
Treat this as a serious risk of outcome-driven analysis and obtain independent methodological scrutiny before proceeding.

Participant exclusions deserve particular care because removing a whole participant may affect several variables and analyses at once. Whether excluding participants becomes falsification rather than a defensible analytical decision depends on the rationale, evidence, consistency, representation, and applicable research practices.

Whatever the decision, do not destroy the original observation simply because it is excluded from a particular analysis. Preserve the raw record and document subsequent changes. This allows reviewers, collaborators, auditors, or future researchers to understand what was changed and why.

07 · A Quick Checklist

Before Removing an Observation, Ask These Questions

Before excluding data from an analysis, check:
What specific scientific, methodological, or data-quality problem makes this observation unsuitable for the analysis?
Is the exclusion criterion prespecified, and if not, why did the need for it emerge later?
Can the reason for exclusion be supported by source records, instrument logs, protocol requirements, or other evidence?
Would you apply the same rule if the observation strongly supported your hypothesis?
Have you applied the criterion consistently to comparable observations?
Have you preserved the original data rather than deleting the evidentiary record?
Where appropriate, have you examined whether the substantive conclusion changes when the observation is retained?
Will your methods and reporting allow readers to understand consequential exclusions and their rationale?
08 · Frequently Asked Questions

Frequently Asked Questions About Removing Research Data

Is removing an outlier falsification?

Not automatically. The researcher needs a defensible reason for the analytical treatment of the outlier. Removing a valid observation merely because it weakens the desired result is very different from excluding a measurement shown to result from instrument failure.

Can I remove data caused by equipment failure?

Potentially, yes. A documented equipment failure can provide a legitimate reason to treat affected measurements as invalid. Preserve the underlying evidence, apply the rule consistently, and report consequential exclusions appropriately.

Can I exclude participants who failed an attention check?

Potentially, but the validity of the attention check, the study design, the exclusion rule, timing of the decision, and consistency of its application matter. An attention check should not become a convenient post hoc device for removing participants with undesirable outcomes.

Is it safer to set an exclusion rule before collecting data?

Prespecification can reduce researcher discretion after outcomes are known and can make the rationale easier to evaluate. It does not guarantee that the criterion is scientifically appropriate, and legitimate unforeseen data-quality problems can still require later decisions.

What if removing one data point changes statistical significance?

That sensitivity is itself useful information. The change does not tell you whether the observation should be excluded, but it increases the importance of a defensible rationale and transparent reporting. Where appropriate, showing analyses with and without the disputed observation can help readers understand the result's robustness.

Should excluded data be deleted from the raw dataset?

Ordinarily, analysis exclusions should not erase the original evidentiary record. Preserve raw data and document how the analytical dataset was derived so that exclusions remain traceable.

Can selective omission of qualitative cases or quotes be falsification?

Potentially, depending on how the selection affects the representation of the research. The same integrity principle can extend beyond numerical datasets, although the methodological standards governing selection differ across research traditions. The more specific issue of selectively showing images, cases, quotes, or examples requires attention to those contextual differences.

09 · The Bottom Line

Removing Data Is Not Falsification Simply Because Data Were Removed

The Bottom Line

Researchers may legitimately exclude data when there is a scientifically defensible reason, but removing or omitting observations can become falsification when it causes the research to be inaccurately represented in the research record.

Do not ask merely whether a data point is inconvenient or unusual. Ask why it should not contribute to the analysis, whether the same rule would be applied regardless of the result, what evidence supports the decision, and whether the exclusion is represented transparently.

10 · Sources and Further Reading

Authoritative Sources on Data Omission and Falsification

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes