03 · What You Need to Know
Good Correction Documentation Preserves Provenance
Start With a Simple Principle: Never Make the Change Mysterious
A correction becomes difficult to evaluate when a reviewer can see that two files differ but cannot determine why.
Good documentation turns an unexplained discrepancy into a traceable research event. It allows someone to move from the original information to the revised information and understand the evidence and reasoning connecting them.
NIH's current intramural research guidance emphasizes that good recordkeeping supports validity, accountability, reproducibility, and research integrity. It also treats procedures for editing, cleaning, and auditing data as part of the research record.
Preserve the Original Information
Documentation begins before the correction itself: retain the evidence showing what was originally recorded.
In an electronic data-capture system, this may occur through an audit trail that preserves prior values. In a computational workflow, the original dataset can remain read-only while cleaning scripts generate derived datasets. In other settings, researchers may retain source documents, dated file versions, laboratory records, or another appropriate record.
The aim is not to keep every accidental temporary file forever. It is to avoid creating a situation in which the corrected value exists but no one can determine what it replaced.
Record the Original and Corrected Values When Applicable
For a discrete data correction, a record such as “Participant 042: examination score changed from 710 to 71” is considerably more informative than “fixed score.”
Not every correction involves replacing one value with another. A duplicate record may be removed from an analytical dataset. A variable may be recoded. A file may be relabeled. An image may be replaced with the correct source image. In those cases, document the original state and resulting state in terms appropriate to the change.
Explain Why the Correction Was Necessary
A correction log should capture the reason for the change rather than merely documenting that a change occurred.
Useful descriptions are specific:
- transcription error from source questionnaire;
- duplicate record created during database import;
- sample identifiers transposed during manual entry;
- incorrect reverse-coding detected in analysis script;
- instrument output associated with wrong specimen identifier.
“Data cleaned” or “value corrected” may be too vague to reconstruct the decision later.
Identify the Evidence Supporting the Correction
The correction should point, where appropriate, to the source that establishes the revised information.
For example, a correction log might identify the original questionnaire, instrument file, laboratory notebook entry, electronic case report form, validated source database, or analysis script used to verify the correction.
This is especially important when the new value is not self-evident. Evidence that a value is wrong does not necessarily establish what the replacement should be.
Correction evidence
Shows why the original record is wrong and, when a replacement is entered, why the revised information is supported.
Correction explanation
Records the nature of the problem and why the documented action was taken.
A robust correction may need both.
Record When the Correction Was Made
The date of correction establishes chronology. This can become important if analyses, reports, submissions, or publications were created before or after the correction.
Do not confuse the date on which the original observation occurred with the date on which its record was corrected.
If a measurement was collected on 5 March but corrected in the database on 20 April, preserving both dates where relevant prevents the correction from masquerading as an original contemporaneous entry.
Record Who Made the Correction When the Research System Requires It
Electronic audit trails commonly associate changes with a user account. Manual correction systems may record initials, signatures, researcher identifiers, or another appropriate attribution mechanism.
This is not merely about assigning blame. Knowing who made a change can help reconstruct the decision, locate supporting information, and establish accountability.
The appropriate level of attribution depends on the research environment, institutional requirements, regulatory obligations, and data-management system.
Do Not Obscure the Original Entry
For paper research records, correction conventions commonly preserve the original information rather than making it unreadable. For electronic records, audit trails can serve a similar function by retaining the previous state.
The broader principle is more important than any particular notation system: correction should add accurate information without rewriting history so completely that the previous record becomes unknowable.
Use an Audit Trail Appropriate to the Research
“Audit trail” does not necessarily mean purchasing specialized clinical-trial software.
Different research environments can preserve change history in different ways. Options include electronic data-capture audit trails, database transaction histories, version-control systems, scripted data-cleaning pipelines, correction logs, file version histories, laboratory information systems, and carefully maintained manual records.
The method should fit the sensitivity, complexity, regulatory requirements, and scale of the research.
Scripted Corrections Can Improve Reproducibility
When researchers work computationally, it is often preferable to preserve the raw dataset and implement systematic corrections in code rather than manually editing hundreds of cells.
A script can show exactly which rule transformed the data and can reproduce the cleaned dataset from the source data. Comments, commit histories, issue records, or associated documentation can explain why the transformation was introduced.
Code does not eliminate the need for justification. A perfectly reproducible unjustified manipulation remains unjustified. Reproducibility shows what happened; scientific documentation explains why it should have happened.
Bulk Corrections Need a Rule, Not Hundreds of Cryptic Notes
Suppose researchers discover that every item on a five-question scale was reverse-coded incorrectly. Recording hundreds of independent corrections may obscure the actual issue.
A clearer record might identify the affected variable set, erroneous coding rule, corrected coding rule, affected dataset versions, script or procedure used to implement the change, and verification performed afterward.
The documentation should match the structure of the error.
Keep Raw, Cleaned, and Analytical Data Conceptually Distinct
A useful workflow often distinguishes the source or raw record from processed data and the final analytical dataset.
This does not mean every project needs three literal spreadsheet files. The implementation may be database-based or generated programmatically. What matters is that researchers can identify which information represents the original record and which information results from subsequent cleaning, correction, derivation, exclusion, or transformation.
This separation makes legitimate data cleaning easier to distinguish from unexplained manipulation.
Document the Consequences of Material Corrections
A correction record should not necessarily stop at the changed cell.
If the corrected information affects derived variables, statistical analyses, figures, tables, participant classifications, or substantive conclusions, those downstream products may need to be regenerated or reviewed.
A useful correction record can therefore include which analyses were rerun and which outputs were replaced.
Published Errors Require a Publication-Level Record
If an error has reached a published article, correcting the private dataset does not by itself correct what readers see.
Researchers should evaluate the effect of the error and follow the journal or publisher's applicable correction process. Depending on the circumstances, the appropriate response may range from a minor correction to more substantial editorial action.
Similarly, datasets deposited in repositories or distributed to collaborators may require updated versions accompanied by version notes or other documentation appropriate to the repository and research context.
Documentation Protects the Researcher as Well as the Research
ORI's current guidance on honest error emphasizes examining research records when distinguishing honest mistakes from potential misconduct. Good records can therefore become important evidence about what actually happened.
Imagine two datasets containing different values for the same participant. Without documentation, the discrepancy invites questions. With a contemporaneous correction record showing the source evidence, date, reason, and resulting reanalysis, the history becomes considerably easier to evaluate.
Watch Out
Do not create a correction trail retrospectively only after someone questions the data and then present it as though it existed at the time of correction. If documentation itself must be reconstructed later, identify that fact accurately and preserve the evidence used for the reconstruction.