Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How Should Legitimate Corrections to Research Data Be Documented?

A legitimate research data correction should leave an intelligible trail from the original record to the corrected one. Document what changed, why it changed, the evidence supporting the correction, when and by whom it was made, and any downstream effects.

458
Documenting Data Corrections Guide 458 of 530
01 · The Question

You Corrected the Data. What Should the Record Show?

Finding and correcting an error is only half of the job. Months or years later, someone may compare two versions of your dataset and discover that Participant 42's value changed from 710 to 71. Without documentation, a legitimate correction and an unexplained alteration can look remarkably similar.

A useful correction record should answer a simple set of questions: what changed, what was there before, why was the change necessary, what evidence supports the new value, when was the correction made, and who made or authorized it?

The exact documentation system varies across research settings, but the underlying objective is stable: preserve provenance so that the corrected record can be understood and verified.

02 · The Short Answer

A Correction Should Be Traceable From the Original Record to the Revised One

In Brief

Document legitimate research data corrections by preserving the original information and recording the corrected value or action, the reason for the correction, the evidence supporting it, the date of the change, and the person or process responsible where appropriate. The correction should be traceable without obscuring the original research record.

For electronic systems, an audit trail may capture much of this information automatically. For other workflows, researchers can use correction logs, version-controlled files, scripted transformations, annotated records, or another method appropriate to the discipline and regulatory context.

03 · What You Need to Know

Good Correction Documentation Preserves Provenance

Start With a Simple Principle: Never Make the Change Mysterious

A correction becomes difficult to evaluate when a reviewer can see that two files differ but cannot determine why.

Good documentation turns an unexplained discrepancy into a traceable research event. It allows someone to move from the original information to the revised information and understand the evidence and reasoning connecting them.

NIH's current intramural research guidance emphasizes that good recordkeeping supports validity, accountability, reproducibility, and research integrity. It also treats procedures for editing, cleaning, and auditing data as part of the research record.

Preserve the Original Information

Documentation begins before the correction itself: retain the evidence showing what was originally recorded.

In an electronic data-capture system, this may occur through an audit trail that preserves prior values. In a computational workflow, the original dataset can remain read-only while cleaning scripts generate derived datasets. In other settings, researchers may retain source documents, dated file versions, laboratory records, or another appropriate record.

The aim is not to keep every accidental temporary file forever. It is to avoid creating a situation in which the corrected value exists but no one can determine what it replaced.

Record the Original and Corrected Values When Applicable

For a discrete data correction, a record such as “Participant 042: examination score changed from 710 to 71” is considerably more informative than “fixed score.”

Not every correction involves replacing one value with another. A duplicate record may be removed from an analytical dataset. A variable may be recoded. A file may be relabeled. An image may be replaced with the correct source image. In those cases, document the original state and resulting state in terms appropriate to the change.

Explain Why the Correction Was Necessary

A correction log should capture the reason for the change rather than merely documenting that a change occurred.

Useful descriptions are specific:

  • transcription error from source questionnaire;
  • duplicate record created during database import;
  • sample identifiers transposed during manual entry;
  • incorrect reverse-coding detected in analysis script;
  • instrument output associated with wrong specimen identifier.

“Data cleaned” or “value corrected” may be too vague to reconstruct the decision later.

Identify the Evidence Supporting the Correction

The correction should point, where appropriate, to the source that establishes the revised information.

For example, a correction log might identify the original questionnaire, instrument file, laboratory notebook entry, electronic case report form, validated source database, or analysis script used to verify the correction.

This is especially important when the new value is not self-evident. Evidence that a value is wrong does not necessarily establish what the replacement should be.

Correction evidence Shows why the original record is wrong and, when a replacement is entered, why the revised information is supported.
Correction explanation Records the nature of the problem and why the documented action was taken.

A robust correction may need both.

Record When the Correction Was Made

The date of correction establishes chronology. This can become important if analyses, reports, submissions, or publications were created before or after the correction.

Do not confuse the date on which the original observation occurred with the date on which its record was corrected.

If a measurement was collected on 5 March but corrected in the database on 20 April, preserving both dates where relevant prevents the correction from masquerading as an original contemporaneous entry.

Record Who Made the Correction When the Research System Requires It

Electronic audit trails commonly associate changes with a user account. Manual correction systems may record initials, signatures, researcher identifiers, or another appropriate attribution mechanism.

This is not merely about assigning blame. Knowing who made a change can help reconstruct the decision, locate supporting information, and establish accountability.

The appropriate level of attribution depends on the research environment, institutional requirements, regulatory obligations, and data-management system.

Do Not Obscure the Original Entry

For paper research records, correction conventions commonly preserve the original information rather than making it unreadable. For electronic records, audit trails can serve a similar function by retaining the previous state.

The broader principle is more important than any particular notation system: correction should add accurate information without rewriting history so completely that the previous record becomes unknowable.

Use an Audit Trail Appropriate to the Research

“Audit trail” does not necessarily mean purchasing specialized clinical-trial software.

Different research environments can preserve change history in different ways. Options include electronic data-capture audit trails, database transaction histories, version-control systems, scripted data-cleaning pipelines, correction logs, file version histories, laboratory information systems, and carefully maintained manual records.

The method should fit the sensitivity, complexity, regulatory requirements, and scale of the research.

Scripted Corrections Can Improve Reproducibility

When researchers work computationally, it is often preferable to preserve the raw dataset and implement systematic corrections in code rather than manually editing hundreds of cells.

A script can show exactly which rule transformed the data and can reproduce the cleaned dataset from the source data. Comments, commit histories, issue records, or associated documentation can explain why the transformation was introduced.

Code does not eliminate the need for justification. A perfectly reproducible unjustified manipulation remains unjustified. Reproducibility shows what happened; scientific documentation explains why it should have happened.

Bulk Corrections Need a Rule, Not Hundreds of Cryptic Notes

Suppose researchers discover that every item on a five-question scale was reverse-coded incorrectly. Recording hundreds of independent corrections may obscure the actual issue.

A clearer record might identify the affected variable set, erroneous coding rule, corrected coding rule, affected dataset versions, script or procedure used to implement the change, and verification performed afterward.

The documentation should match the structure of the error.

Keep Raw, Cleaned, and Analytical Data Conceptually Distinct

A useful workflow often distinguishes the source or raw record from processed data and the final analytical dataset.

This does not mean every project needs three literal spreadsheet files. The implementation may be database-based or generated programmatically. What matters is that researchers can identify which information represents the original record and which information results from subsequent cleaning, correction, derivation, exclusion, or transformation.

This separation makes legitimate data cleaning easier to distinguish from unexplained manipulation.

Document the Consequences of Material Corrections

A correction record should not necessarily stop at the changed cell.

If the corrected information affects derived variables, statistical analyses, figures, tables, participant classifications, or substantive conclusions, those downstream products may need to be regenerated or reviewed.

A useful correction record can therefore include which analyses were rerun and which outputs were replaced.

Published Errors Require a Publication-Level Record

If an error has reached a published article, correcting the private dataset does not by itself correct what readers see.

Researchers should evaluate the effect of the error and follow the journal or publisher's applicable correction process. Depending on the circumstances, the appropriate response may range from a minor correction to more substantial editorial action.

Similarly, datasets deposited in repositories or distributed to collaborators may require updated versions accompanied by version notes or other documentation appropriate to the repository and research context.

Documentation Protects the Researcher as Well as the Research

ORI's current guidance on honest error emphasizes examining research records when distinguishing honest mistakes from potential misconduct. Good records can therefore become important evidence about what actually happened.

Imagine two datasets containing different values for the same participant. Without documentation, the discrepancy invites questions. With a contemporaneous correction record showing the source evidence, date, reason, and resulting reanalysis, the history becomes considerably easier to evaluate.

Watch Out

Do not create a correction trail retrospectively only after someone questions the data and then present it as though it existed at the time of correction. If documentation itself must be reconstructed later, identify that fact accurately and preserve the evidence used for the reconstruction.

04 · A Practical Example

What a Useful Data Correction Record Looks Like

Hypothetical Example

Correcting a Transcription Error

A researcher discovers that Participant P042's outcome score was entered incorrectly during manual transcription.

Correction field Example documentation
Affected record Participant P042, post-test score
Original dataset value 710
Corrected value 71
Reason Manual transcription error
Supporting evidence Original scored assessment and electronic assessment record both show 71
Correction date 4 October 2026
Action Analytical dataset corrected; original imported dataset retained
Downstream check Affected descriptive statistics and regression analysis rerun

The precise fields needed in a real project will vary. The important feature is that another qualified person can understand the correction without having to infer why 710 mysteriously became 71.

05 · What Researchers Often Get Wrong

Common Mistakes When Documenting Data Corrections

Misconception

Keeping the Final Correct Dataset Is Enough

The final dataset may contain the right values while providing no evidence of how they got there. Preserving appropriate source records and change history makes consequential corrections verifiable rather than merely asserted.

Misconception

“Corrected Error” Is a Sufficient Explanation

Usually not for a consequential correction. Document what kind of error occurred and what evidence supported the correction. Specificity makes the record useful months or years later.

Misconception

An Audit Trail Means I Do Not Need to Explain Anything

An automated audit trail may show that a value changed, who changed it, and when. It may not explain the scientific or evidentiary reason for the change. Technical traceability and methodological justification serve related but different purposes.

Misconception

Every Typo Requires an Elaborate Formal Investigation

Documentation should be proportionate to the research setting and significance of the correction. A minor obvious typo and a correction that reverses the study's primary conclusion need not receive identical treatment. Applicable institutional or regulatory requirements still take precedence.

Misconception

If the Correction Weakens the Result, Documentation Matters Less

The direction of the effect is irrelevant to good provenance. Corrections should be documented because the research record changed, not only when the change appears favorable or potentially suspicious.

06 · What This Means for You

Build a Correction Record That Another Researcher Could Reconstruct

You do not need a bureaucratic ceremony every time a typo is found. You do need a correction system proportionate to the research that makes consequential changes intelligible.

A practical documentation framework

If one discrete value is corrected
Preserve the original, record the corrected value, reason, supporting source, date, and responsible person or process as appropriate.
If a systematic error affects many records
Document the error mechanism, affected records or variables, correction rule, implementation method, and verification rather than relying on hundreds of unexplained manual edits.
If the correction is implemented computationally
Preserve the source data and code or processing record that reproduces the corrected dataset.
If the correction changes an analysis or conclusion
Record which analyses and outputs were rerun or revised and determine whether additional disclosure is required.
If the erroneous information has already been shared or published
Correct the relevant external research record through the applicable repository, journal, institutional, or other process.

The documentation should ultimately support the principle behind correcting genuine data errors without creating a misleading record: the correction should make the research more accurate while leaving enough evidence to show why the change was legitimate.

Good versioning and provenance also contribute to the broader task of preserving evidence that research data and images are authentic.

07 · A Quick Checklist

What to Record When Research Data Are Corrected

For a consequential data correction, document:
The dataset, record, variable, image, file, or other research material affected by the correction.
The original value or state, preserved through the source record, raw data, audit trail, or another appropriate mechanism.
The corrected value or action taken.
The specific reason the original information was determined to be erroneous.
The source evidence or reproducible rule supporting the correction.
The date of the correction and the responsible researcher, user, or process where appropriate.
Any systematic rule applied to other records affected by the same error.
Which analyses, tables, figures, reports, datasets, or other downstream outputs were regenerated or reviewed.
Any additional correction required for data already submitted, shared, deposited, presented, or published.
08 · Frequently Asked Questions

Frequently Asked Questions About Documenting Data Corrections

Do I need a correction log for every research project?

The appropriate system depends on the type of research, institutional requirements, applicable regulations, and complexity of the data. The underlying need is traceability. Consequential corrections should not become unexplained differences between dataset versions.

What should a data correction log contain?

Useful fields commonly include the affected record, original value or state, corrected value or action, reason for the correction, supporting evidence, correction date, and responsible person or process where appropriate. Material downstream effects may also need documentation.

Should I edit the raw data file directly?

Where practical, preserve the source or raw record and implement corrections in a traceable working or analytical dataset, database process, or scripted pipeline. Specific requirements vary by research environment and regulatory context.

Is version control enough to document corrections?

Version control can provide an excellent technical history of what changed, particularly for code and text-based data. It does not automatically explain why a scientific correction was justified, so meaningful commit messages, issue records, correction logs, or related documentation may still be needed.

What if an electronic system already has an audit trail?

Use the system's audit-trail functions rather than bypassing them. Check whether the trail captures the information required in your setting, including previous values, timestamps, users, and reasons for changes where applicable.

How should I document hundreds of corrections caused by one coding error?

Document the systematic error and correction rule clearly, identify the affected variables or records, preserve the code or procedure used to implement the correction, and verify the resulting dataset. A reproducible rule may be more informative than hundreds of disconnected manual notes.

Do I need to report data corrections in the published paper?

Not every routine internal correction requires a detailed narrative in a publication. Reporting should nevertheless describe consequential data-processing procedures sufficiently for the research to be understood, and material errors discovered after dissemination may require formal correction under the relevant journal or publisher policy.

Can good correction records help if someone alleges manipulation?

Yes. Contemporaneous records showing the original information, evidence supporting the correction, timing, rationale, and resulting analyses can help establish what actually occurred. They do not automatically resolve every integrity question, but they provide considerably stronger evidence than an unexplained altered dataset.

09 · The Bottom Line

A Good Correction Leaves Evidence of Both the Error and the Fix

The Bottom Line

Legitimate research data corrections should be documented so that another qualified person can determine what changed, what the original information was, why the correction was necessary, what evidence supported it, and when and by whom the change was made where appropriate.

Preserve provenance rather than merely preserving the final clean dataset. Whether you use an electronic audit trail, scripted workflow, version history, correction log, or another appropriate system, the objective is the same: a corrected research record whose history remains intelligible.

10 · Sources and Further Reading

Authoritative Sources on Research Records and Data Corrections

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes