03 · What You Need to Know
The Problem Is Not Removal Itself but What the Removal Does to the Research Record
Falsification Explicitly Includes Certain Omissions
Under the U.S. Public Health Service definition, falsification includes manipulating research materials, equipment, or processes, or changing or omitting data or results such that the research is not accurately represented in the research record.
The word “omitting” matters. A researcher does not need to change a numerical value to falsify a record. Selectively removing genuine observations can potentially create an inaccurate representation of the research.
But the definition does not say that every omission is falsification.
Researchers Are Sometimes Supposed to Exclude Data
Scientific analyses often require eligibility criteria, quality-control rules, handling of measurement failures, treatment of duplicate records, and decisions about observations that cannot validly contribute to a particular analysis.
The federal research misconduct policy explicitly recognized this distinction when responding to concerns about omission. It explains that appropriate omission can occur under accepted research practices, while omission becomes falsification when it misleads the reader about the research results.
That distinction prevents an absurd interpretation in which ordinary quality control becomes misconduct merely because the analyzed dataset contains fewer rows than the raw file.
“Bad Data” Is Too Vague to Justify Anything
Calling an observation “bad” tells you almost nothing scientifically.
Was the instrument outside its calibrated range? Did the participant violate an eligibility criterion? Was the observation duplicated because of an import error? Did the researcher establish that the value resulted from contamination? Or does “bad” merely mean that the observation weakens the hypothesis?
Invalid or unusable observation
There is a scientifically defensible reason why the observation should not contribute to the specified analysis.
Inconvenient observation
The observation is valid but undesirable because it weakens, complicates, or contradicts the preferred result.
Those categories are not always easy to determine, but researchers should not collapse them into the same label.
Outlier Status Alone Does Not Make a Data Point Invalid
An outlier is unusual relative to other observations. It is not necessarily erroneous.
A very high value might reflect a data-entry error. It might also represent genuine biological variation, an unusual participant, an unexpected social phenomenon, a rare but legitimate event, or evidence that the assumed statistical model fits poorly.
Removing an observation solely because it is extreme can therefore discard scientifically meaningful information. Researchers should investigate why the observation is unusual before deciding what analytical treatment is appropriate.
Prespecified Exclusion Rules Are Usually Easier to Defend
When exclusion criteria are established before researchers know which observations support their preferred result, there is less opportunity for the observed outcome to drive the decision.
Examples might include excluding participants who fail predefined eligibility requirements, measurements produced during documented equipment failure, or trials invalidated according to established quality-control criteria.
Prespecification does not make every rule scientifically sound. Nor does the absence of prespecification automatically make every later exclusion improper. Unexpected problems arise during research. But post hoc exclusions deserve particularly careful justification because the researcher already knows what the data look like.
A Post Hoc Exclusion Can Be Legitimate
Suppose researchers discover during analysis that a software bug duplicated several observations. The exclusion rule could not have been specified before the bug was discovered. Removing the duplicates may nevertheless make the dataset more accurate.
Similarly, an instrument malfunction may become apparent only after data collection. The relevant issue is whether there is evidence supporting the exclusion and whether the rule is applied consistently rather than opportunistically.
The broader boundary between legitimate processing and data manipulation during cleaning therefore cannot be reduced to “before analysis good, after analysis bad.”
Outcome-Driven Exclusion Is a Serious Warning Sign
Consider a researcher who runs an analysis, dislikes the result, removes one participant, reruns the model, removes another, and continues until the desired effect becomes statistically significant.
That sequence is fundamentally different from discovering that a sensor failed for a documented reason.
The problem becomes particularly serious if the exclusions are then concealed and the final dataset or report is presented as though those observations were never part of the research. Depending on the facts, such omission may cause the research to be inaccurately represented and therefore fall within falsification.
The Same Exclusion Rule Should Normally Apply to Comparable Cases
Selective enforcement can be revealing.
Suppose a researcher claims that scores more than three standard deviations from a specified reference point are invalid. If the rule removes only observations unfavorable to the hypothesis while equally extreme favorable observations remain, the stated rationale may not explain the actual decision process.
Consistency does not by itself prove that a rule is scientifically appropriate, but inconsistent application should prompt scrutiny.
Transparency Does Not Automatically Rescue a Scientifically Indefensible Exclusion
Reporting an exclusion is better than hiding it, but disclosure alone does not make the analytical choice sound.
A paper could openly state that five observations were removed because they prevented statistical significance. Transparency would allow readers to see the problem, but it would not convert the exclusion into a defensible methodological decision.
Researchers therefore need both scientific justification and accurate reporting.
Sometimes the Best Answer Is to Show the Analysis Both Ways
When a questionable observation has a plausible but uncertain reason for exclusion, a sensitivity analysis can be informative. Researchers may compare results with and without the observation and explain why one analysis is primary.
This does not solve every exclusion problem, but it prevents an ambiguous data point from silently determining the story. If the substantive conclusion changes dramatically depending on one defensible analytical decision, readers may need to know that.
Formal Falsification Requires More Than a Debatable Analytical Choice
A questionable exclusion does not automatically establish research misconduct. Under the PHS framework, misconduct excludes honest error and differences of opinion, and a finding requires a significant departure from accepted practices, the required culpable state of mind, and proof by a preponderance of the evidence.
Methodological criticism and misconduct findings are therefore not interchangeable. Researchers can make poor analytical choices without necessarily committing falsification. Conversely, describing an exclusion as an “analytical decision” does not shield deliberate or reckless manipulation that makes the research record inaccurate.
Watch Out
If your reason for removing an observation begins with “the result becomes better without it,” stop. The effect of an exclusion on your preferred conclusion is not, by itself, a scientific reason to exclude the observation.
06 · What This Means for You
Ask Why the Observation Is Invalid Before Asking Whether You Want It
The strongest exclusion decisions begin with a scientific or methodological reason, not with the direction of the result.
A simple decision framework
If an observation meets a valid prespecified exclusion criterion
Apply the criterion consistently and document the exclusion according to the study's procedures.
If a previously unknown data-quality problem is discovered
Investigate the source, formulate a scientifically defensible rule, apply it consistently, preserve the original data, and document the decision.
If an observation is merely extreme or inconvenient
Do not assume that unusual means invalid. Investigate it and consider robust methods or sensitivity analyses where appropriate.
If the exclusion was considered only after seeing that it improves the desired result
Treat this as a serious risk of outcome-driven analysis and obtain independent methodological scrutiny before proceeding.
Participant exclusions deserve particular care because removing a whole participant may affect several variables and analyses at once. Whether excluding participants becomes falsification rather than a defensible analytical decision depends on the rationale, evidence, consistency, representation, and applicable research practices.
Whatever the decision, do not destroy the original observation simply because it is excluded from a particular analysis. Preserve the raw record and document subsequent changes. This allows reviewers, collaborators, auditors, or future researchers to understand what was changed and why.
07 · A Quick Checklist
Before Removing an Observation, Ask These Questions
Before excluding data from an analysis, check:
What specific scientific, methodological, or data-quality problem makes this observation unsuitable for the analysis?
Is the exclusion criterion prespecified, and if not, why did the need for it emerge later?
Can the reason for exclusion be supported by source records, instrument logs, protocol requirements, or other evidence?
Would you apply the same rule if the observation strongly supported your hypothesis?
Have you applied the criterion consistently to comparable observations?
Have you preserved the original data rather than deleting the evidentiary record?
Where appropriate, have you examined whether the substantive conclusion changes when the observation is retained?
Will your methods and reporting allow readers to understand consequential exclusions and their rationale?