03 · What You Need to Know
A Correction Should Make the Research Record More Accurate, Not More Convenient
Research Misconduct Does Not Include Honest Error
The U.S. Public Health Service definition of research misconduct expressly excludes honest error and differences of opinion. ORI's current guidance on honest error further recognizes that actions that initially appear to involve fabrication or falsification may, after examination of the evidence, turn out to be honest mistakes.
That distinction matters because real research generates errors. Numbers are mistyped. Files are imported incorrectly. Labels are transposed. Formulas reference the wrong cells. Software code contains bugs. Human beings remain stubbornly involved in research, despite several centuries of methodological refinement.
Discovering an error therefore does not establish misconduct. The next question is what the researcher does after discovering it.
Falsification Is About Inaccurate Representation, Not Every Change to Data
Under the PHS definition, falsification includes manipulating research materials, equipment, or processes, or changing or omitting data or results such that the research is not accurately represented in the research record.
The final phrase is crucial.
If a source questionnaire records an age of 23 but a transcription mistake produced 223 in the electronic dataset, changing 223 back to 23 makes the dataset more accurately represent the underlying record. The correction changes the data file, but the purpose and effect of the change are opposite to falsification.
Genuine correction
Reliable evidence shows that the existing record is erroneous and supports the corrected information.
Improper alteration
Data are changed without adequate evidentiary support or are selectively altered so that the research record presents a misleading account.
A Suspicious Value and a Known Error Are Not the Same Thing
Before correcting anything, establish what you actually know.
A value of 223 for age may be obviously impossible in a study of living human participants. That establishes that something is wrong. It does not necessarily establish that the correct value is 23.
Perhaps 223 represents a coding error, a shifted decimal, a participant identifier entered in the wrong field, or some other problem. If the original questionnaire or verified source record shows 23, the correction has an evidentiary basis. Without that evidence, replacing 223 with whichever plausible value seems most likely risks creating information that was never verified.
The appropriate treatment of an unresolved error may therefore be to flag the value, mark it missing, exclude it from a specific analysis when justified, or investigate further rather than guess.
Correct From Evidence, Not From What Makes the Dataset Look Better
A defensible correction should be linked to a reliable source or reproducible rule.
Depending on the research, that evidence might include an original questionnaire, laboratory notebook, instrument output, source document, database record, image metadata, electronic audit trail, validated calculation, analysis script, or another contemporaneous record.
The principle is the same across very different research settings: the corrected value should have a reason for being there beyond “this looks more sensible.”
The Original Record Should Usually Remain Recoverable
Good correction practices do not make the earlier record disappear without explanation.
NIH guidance emphasizes that good scientific recordkeeping is essential to validity, accountability, reproducibility, and research integrity. Research records encompass not only data but also procedures for editing, cleaning, auditing, and otherwise managing them.
For electronic systems, an audit trail may automatically retain previous values, timestamps, users, and reasons for changes. In other environments, researchers may maintain raw data separately from cleaned data, use version control, retain correction logs, or document amendments through another appropriate mechanism.
The implementation varies. The principle does not: someone reviewing a consequential correction later should be able to determine what happened.
Do Not Overwrite Raw Data Merely to Make Them “Correct”
Suppose the raw instrument file contains an erroneous value because the instrument malfunctioned. Researchers may appropriately decide that the value should not enter the primary analysis. That does not mean the original instrument output should be edited until the malfunction disappears from history.
Raw or source records often serve as evidence of what was originally recorded. Analytical datasets can contain documented corrections or derived values while preserving the source from which those decisions were made.
This distinction is particularly important when the correction itself might later be questioned.
A Correction Can Change the Result and Still Be Legitimate
Researchers sometimes worry that correcting an error will look suspicious if it changes statistical significance, reverses a comparison, or otherwise affects the conclusion.
The effect of the correction does not determine whether the correction is legitimate.
If reliable evidence establishes that a value is wrong, the error should not be preserved merely because correcting it changes the result. Conversely, a correction does not become valid merely because it improves the result.
The evidence supporting the correction should drive the decision.
Apply Systematic Corrections Systematically
Some errors affect more than one observation. Perhaps a coding script accidentally reversed every response on a particular scale, or a laboratory instrument applied an incorrect calibration factor during a defined period.
Once the problem is established, the correction rule should ordinarily be applied to all affected records rather than only to observations whose correction helps the hypothesis.
Selective correction can transform a legitimate quality-control procedure into something much more difficult to defend.
Correcting Data Is Different From Filling Gaps With Memory
A verified correction has evidence supporting both propositions: the existing entry is wrong, and the replacement is the correct value.
A missing value reconstructed only from recollection may not have that evidentiary support. This is why filling in a missing value from memory requires greater caution. Knowing that something is missing does not automatically establish what belongs in the gap.
Corrections Made After Seeing the Results Are Not Automatically Improper
Many errors are discovered during analysis precisely because researchers inspect distributions, run diagnostics, compare records, or encounter implausible results.
Finding the error after seeing outcomes does not prohibit correction. It does increase the value of documenting objective evidence for the change, particularly if the correction materially affects the conclusion.
Researchers should be able to show that the correction would have been made regardless of whether it strengthened or weakened the preferred finding.
Correcting Published or Submitted Results May Require More Than Updating Your Dataset
If an error has already propagated into a manuscript, report, repository, conference presentation, thesis, or published article, correcting the local dataset may not fully correct the research record.
The appropriate response depends on where the erroneous information has been disseminated and how materially it affects the work. Researchers may need to notify collaborators, supervisors, repositories, journals, institutions, funders, or other relevant parties according to applicable policies and circumstances.
The goal is not merely to possess a correct private file. It is to ensure that consequential versions of the research record are not left materially inaccurate.
Watch Out
“I know this value is wrong” and “I know the correct value” are two different claims. Do not turn evidence of an error into permission to invent its replacement. If the correct value cannot be established, preserve that uncertainty.