Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Should Researchers Preserve to Demonstrate That Their Data and Images Are Authentic?

Researchers should preserve enough original data, metadata, processing history, code, and contextual records to trace reported results back to their sources. Authenticity is much easier to demonstrate when provenance survives alongside the final figure or dataset.

464
Preserving Authentic Research Data Guide 464 of 530
01 · The Question

If Someone Questions Your Data Years Later, What Would You Need to Show Them?

A journal asks for the uncropped Western blot behind a published panel. A collaborator wants to verify where a microscopy image came from. A reviewer questions a corrected data point. An institution needs to determine whether two apparently duplicated images actually represent different experiments.

The final manuscript is not enough to answer these questions.

Researchers need records connecting published or analyzed material back to the research that produced it. Authenticity is demonstrated through provenance: evidence of where data came from, what happened to them, and how the final result was derived.

02 · The Short Answer

Preserve the Evidence Needed to Reconstruct the Research

In Brief

Researchers should preserve original or primary data and images, relevant metadata and source records, protocols and laboratory records, processing and cleaning history, analysis code, correction and audit trails, and the information needed to connect final figures and results to their underlying sources. The precise records and retention period depend on the research type, institution, funder, jurisdiction, agreements, and applicable regulations.

For image-based research, unprocessed image files are particularly important because image authentication may require access to the original data. Preserve provenance before publication rather than trying to reconstruct it only after a question arises.

03 · What You Need to Know

Authenticity Depends on More Than Keeping the Final Dataset

Preserve the Original or Primary Research Data

The analytical dataset is not necessarily the primary research record.

Depending on the study, primary data may include instrument outputs, survey responses, experimental measurements, observational records, interview recordings or transcripts where appropriate, laboratory notebook entries, clinical research records, photographs, microscopy images, electrophoretic gels, sensor files, field notes, or other direct records of the research.

NIH's current intramural research guidance states that primary data should be retained, including observations and experiments that do not directly lead to publication. It specifically notes that usable confocal microscopy imaging files should generally be retained in their original format, subject to its stated qualifications for technically problematic or exceptionally large imaging data.

The general principle is powerful: do not preserve only the evidence that eventually appears in the paper.

Keep Unprocessed Image Files

For scientific images, the original unprocessed files can be essential.

ORI's forensic-image guidance states that authenticating a scientific image requires access to original data. An image discrepancy in a published figure can identify a question, but determining what happened often requires comparison with the source image.

Nature Portfolio similarly advises researchers to retain unprocessed data and metadata after publication and notes that editors may request unprocessed files during peer review or when investigating post-publication concerns.

If the only surviving version is a cropped, compressed, contrast-adjusted JPEG copied from a PowerPoint slide, much of the evidence needed to establish provenance may already be gone.

Preserve Native Formats and Relevant Metadata

Converting every file into a convenient generic format can discard information.

Native instrument or image formats may retain acquisition settings, timestamps, dimensions, channel information, calibration data, bit depth, device information, or other metadata that do not survive ordinary export.

NIH record-retention guidance for clinical research also advises keeping electronic records in native format where applicable. The specific requirement varies by research environment, but preserving native files can materially strengthen provenance.

For image research, metadata may help establish when and how an image was acquired. For computational data, metadata may define variables, units, coding schemes, versions, or provenance relationships.

Preserve Enough Context to Understand What Each File Represents

A folder containing 10,000 files named IMG001.tif through IMG10000.tif may technically preserve the images while making their scientific provenance nearly impossible to reconstruct.

Researchers need enough contextual information to determine which sample, participant, specimen, condition, time point, experiment, replicate, instrument, or procedure produced each relevant record.

This can come from structured file naming, laboratory notebooks, electronic laboratory information systems, manifests, database identifiers, metadata, sample maps, study logs, or other systems appropriate to the research.

The objective is that another qualified person should be able to connect the file to the research event it represents.

Preserve Acquisition Information When Interpretation Depends on It

For image-based research, the acquisition process can matter as much as subsequent processing.

Microscope settings, exposure times, detector parameters, objective information, acquisition software, channel settings, instrument calibration, and related information may be necessary to understand whether images are genuinely comparable.

Nature Portfolio image-integrity guidance asks authors to list image-acquisition tools and processing software and document key acquisition settings and processing manipulations in the Methods.

Not every parameter needs indefinite preservation in every project. Preserve what is necessary to interpret, reproduce, and authenticate the evidence according to the standards of the field.

Keep the Processing History Between the Original and Final Image

An original image and final figure are two ends of a chain. Researchers should be able to explain what happened between them.

Relevant records may include crop boundaries, brightness and contrast adjustments, channel combinations, pseudocoloring, nonlinear transformations, stitching, segmentation, thresholding, annotations, lane rearrangements, and figure assembly.

The amount of documentation needed depends on the operation and discipline. Routine processing may be captured in a reproducible script or standard workflow. Consequential manual edits may require more explicit documentation.

The aim is to make image processing distinguishable from unexplained image alteration.

Preserve the Full Image When Only a Crop Is Published

A crop cannot demonstrate what existed outside its boundaries.

If a paper publishes only part of a microscopy field, gel, blot, radiological image, photograph, or other image, preserve the corresponding uncropped original.

This can establish whether important information was omitted and whether the published region came from the claimed source. Nature Portfolio requires unprocessed original gel and Western blot images for accepted life-science papers and recommends retaining unprocessed data and metadata after publication.

Keep Figure-to-Source Provenance

Final figures frequently contain many panels. Months later, even their creators may struggle to remember which original file produced Panel 4C.

Maintain a mapping between figure panels and source files where practical. This might be a figure manifest, reproducible script, notebook, spreadsheet, file-naming convention, electronic lab notebook entry, or another system.

Source file The original data or image from which the displayed evidence derives.
Figure provenance The traceable relationship connecting the source file, processing steps, selected region or measurement, and final published panel.

This becomes particularly important when figures combine images from different experiments or files.

Preserve Raw Data Separately From Cleaned or Analytical Data

Researchers often correct errors, remove duplicates, recode variables, derive measures, and create analysis-ready datasets. Those operations need not threaten authenticity if the original record remains available and the transformations are traceable.

A useful workflow preserves source data separately and derives cleaned datasets through documented procedures. Depending on the project, this might involve scripts, database audit trails, versioned datasets, correction logs, or another reproducible process.

Overwriting the only copy of the original dataset makes later verification much harder.

Keep Data Dictionaries, Codebooks, and Variable Definitions

A numerical dataset without definitions may be impossible to interpret correctly.

Preserve information explaining variable names, units, missing-value codes, category definitions, derived variables, scoring rules, transformations, and other conventions needed to understand the data.

For qualitative research, analogous documentation may include coding frameworks, methodological memos, code definitions, and records connecting coded excerpts to source material, subject to ethical and confidentiality constraints.

Preserve Analysis and Processing Code

When results depend on software scripts, code is part of the evidentiary chain.

Preserving the exact or appropriately versioned code used to clean, transform, analyze, and visualize data can show how reported results were produced from the underlying records.

Dependencies, software versions, random seeds, configuration files, computational environments, or workflow definitions may also matter for complex analyses.

Code alone is not enough if the underlying data or documentation are missing, but it can make the path from source data to reported result substantially more transparent.

Preserve Correction and Audit Trails

Datasets legitimately change. A value may be corrected from a source document, a duplicate removed, a coding error fixed, or an identifier reconciled.

Keep enough information to show what changed and why. Depending on the system, this may include previous values, corrected values, timestamps, users, reasons for changes, source evidence, or systematic correction rules.

The principles for documenting legitimate research data corrections help distinguish a traceable correction from an unexplained alteration.

Preserve Excluded and Unpublished Data Where Required

Do not assume that only observations appearing in the final analysis matter.

Excluded observations can be important for evaluating whether exclusion criteria were applied appropriately. Additional experimental replicates can show whether a displayed image was genuinely representative. Unsuccessful or negative experiments may be relevant to reconstructing the research history.

NIH intramural guidance explicitly states that primary data from observations and experiments not directly leading to publication must also be retained under its policy.

The precise retention obligation elsewhere depends on the applicable institutional, funder, regulatory, contractual, and disciplinary requirements.

Preserve Versions of Manuscripts, Figures, and Analytical Outputs When They Matter

Research provenance extends beyond raw data.

Draft figures may show when panels were replaced or rearranged. Analysis outputs may document which model produced a reported table. Manuscript versions can establish how interpretations changed after corrections or review.

Researchers do not necessarily need to archive every autosaved draft forever. The principle is proportionality: preserve records needed to reconstruct consequential decisions and reported results.

Use Backups That Protect Against Both Loss and Silent Alteration

A file existing in one location is not a preservation strategy.

NIH intramural guidance recommends regular backup of electronic research records and archival storage designed to prevent subsequent alteration. Secure institutional storage, access controls, backup systems, checksums, immutable archives, version histories, or other mechanisms can help protect records depending on the research environment.

Storage also needs to respect participant confidentiality, security requirements, controlled-access restrictions, intellectual property, and other legal or ethical obligations. “Keep everything everywhere” is not a responsible data-management policy.

Content Provenance Is Becoming More Important as Image Generation Improves

ORI now specifically highlights content provenance as an emerging research-integrity issue in response to increasingly powerful automated and AI-enabled image manipulation. It defines content provenance as a verifiable record of where digital content originated, modifications made to it, and where it has been published.

That concept is likely to become increasingly relevant as visual evidence becomes easier to synthesize or alter convincingly.

Traditional research records remain essential, but cryptographic or machine-verifiable provenance systems may eventually supplement laboratory notebooks, metadata, audit trails, and source-file archives in some research settings.

There Is No Universal Retention Period for Every Research Record

Researchers should be cautious about memorizing one number and treating it as universal.

Retention periods vary by institution, funder, award, jurisdiction, research type, regulatory framework, intellectual-property considerations, participant protections, and whether an audit, claim, investigation, or other proceeding is ongoing.

For example, current NIH intramural guidance generally requires its intramural research records to be retained for seven years after project completion or until no longer needed for scientific reference, whichever is longer, with longer periods for some records. NIH grant recipients operate under different grant record-retention requirements, including exceptions and qualifications.

Researchers should therefore verify the policy that actually governs their project rather than importing a retention period from another institution or research context.

Watch Out

Do not wait for a journal query or integrity allegation before trying to reconstruct provenance. Once original files, metadata, processing histories, or excluded observations are lost, later explanations may be impossible to verify even when the research was entirely legitimate.

04 · A Practical Example

What Would You Need to Authenticate a Published Microscopy Figure?

Hypothetical Example

A Journal Questions Two Similar Figure Panels

Three years after publication, a reader notices that two microscopy panels look unusually similar. The journal asks the authors to explain their provenance.

Original image archive The laboratory retains the native unprocessed microscopy files, including acquisition metadata and timestamps.
Experimental records The laboratory notebook and sample manifest identify which specimen, condition, replicate, and acquisition session correspond to each file.
Figure provenance A figure map identifies the exact source file used for each published panel and the crop coordinates or region selected.
Processing history The researchers retain the workflow showing the global display adjustments, channel assignments, scale-bar generation, and figure assembly.
Resolution The source records show that the panels came from two distinct images with similar biological structures rather than one duplicated image.

The researchers do not have to rely on “we remember that they were different experiments.” Their records allow the claim to be checked.

05 · What Researchers Often Get Wrong

Common Mistakes in Preserving Data and Image Authenticity

Misconception

The Published Figure Is Enough to Preserve the Evidence

A publication-ready figure may be cropped, resized, compressed, annotated, assembled, or otherwise processed. It cannot necessarily establish the original data or the provenance of individual panels. Preserve the underlying source files.

Misconception

I Only Need to Keep Data That Appeared in the Paper

Applicable policies may require much broader retention, and excluded or unpublished observations can be important for understanding how the reported evidence was selected. Verify the requirements governing your research rather than retaining only the final success stories.

Misconception

A JPEG Export Is the Same as the Original Image

Export and compression can remove metadata, reduce bit depth, alter pixel information, or introduce artifacts. Preserve native or original acquisition files where required and practical rather than relying only on presentation copies.

Misconception

If My Analysis Code Is Saved, My Research Is Reproducible

Code cannot reproduce an analysis if the required data, variable definitions, software environment, dependencies, or processing information are missing. Preserve the relevant components of the entire provenance chain.

Misconception

There Is One Standard Number of Years That All Research Data Must Be Kept

No. Retention requirements vary substantially. Institutional, funder, regulatory, contractual, intellectual-property, and disciplinary requirements can all affect how long particular records must be preserved.

06 · What This Means for You

Preserve a Chain From Research Event to Reported Result

A useful preservation strategy works backward from a simple question: if someone challenged this result later, what would they need to verify where it came from?

A simple provenance framework

If a result originates from an instrument, image, questionnaire, observation, or other primary source
Preserve the original or primary record and enough metadata to identify what it represents.
If the original data are cleaned, corrected, transformed, or recoded
Preserve the source version and a traceable record of the transformations leading to the analytical data.
If a figure contains selected, cropped, or processed images
Retain the unprocessed sources and enough information to map each displayed panel back to its origin.
If results are generated computationally
Preserve the code, parameters, relevant software or environment information, and data needed to reconstruct the reported output.
If you are deciding how long to retain records
Check the requirements of your institution, funder, award, regulator, contracts, journal, and jurisdiction rather than assuming a universal period.

Make provenance part of the workflow rather than an archaeological project performed after publication. A well-organized record should allow the path from source evidence to cleaned data, analysis, figure, and reported conclusion to be reconstructed without relying primarily on personal memory.

07 · A Quick Checklist

What to Preserve to Support Data and Image Authenticity

For research records that support reported findings, consider preserving:
Original or primary data, including relevant observations and experiments that did not ultimately appear in the publication, according to applicable requirements.
Native, unprocessed image files and relevant acquisition metadata.
Laboratory notebooks, source documents, protocols, sample maps, study logs, or equivalent records connecting files to research events.
Data dictionaries, codebooks, units, missing-value definitions, scoring rules, and other information needed to interpret the data.
Cleaning, correction, exclusion, transformation, and audit histories needed to trace analytical data back to their sources.
Analysis and image-processing code, parameters, software information, and reproducible workflows where applicable.
Mappings between final figure panels and their original source images or datasets.
Relevant versions of analytical outputs, figures, and other records needed to reconstruct consequential research decisions.
Secure backups and archival copies appropriate to confidentiality, access, integrity, and retention requirements.
The current retention requirements that specifically govern your project rather than a generic retention period borrowed from another setting.
08 · Frequently Asked Questions

Frequently Asked Questions About Preserving Authentic Research Records

Should I keep every original research image?

Follow the retention requirements and accepted practices applicable to your project. NIH intramural guidance, for example, requires retention of primary data and specifically addresses original-format confocal microscopy files, while recognizing limited exceptions for technically problematic or prohibitively large imaging data. Do not assume that only images selected for publication need preservation.

Should I keep RAW or native image files after exporting TIFFs or JPEGs?

Where native files constitute the original research record, preserve them according to applicable requirements. Exported files may lose metadata or other information useful for interpretation and authentication. Presentation copies should not automatically replace source data.

Do I need to keep images that did not support my hypothesis?

Applicable record-retention requirements may extend to primary data regardless of whether the observation was published or supported the hypothesis. Such images may also be important for demonstrating whether published examples were representative.

How long should research data be retained?

There is no universal retention period. Requirements depend on the institution, funder, award, jurisdiction, regulatory framework, contracts, intellectual-property issues, research type, and whether audits, investigations, claims, or other proceedings are pending. Verify the rules governing the specific project.

Is cloud storage enough for preserving research data?

Only if the storage arrangement satisfies the project's institutional, security, confidentiality, access, backup, integrity, and retention requirements. Consumer cloud storage should not automatically be assumed appropriate for sensitive or regulated research records.

What if the original data are already missing?

Do not recreate source data merely to make the record appear complete. Preserve what remains, document the loss, recover information from reliable surviving sources where possible, and seek appropriate institutional guidance when the missing records are consequential. Any retrospective reconstruction should be identified as such.

Why are original files important if the published figure looks normal?

ORI notes that authenticating a scientific image requires access to original data. A published figure may be too processed, compressed, cropped, or assembled to establish its provenance. Original files allow questioned features to be compared with the underlying evidence.

Should I preserve AI provenance information if AI tools were used on research images?

Preserve records sufficient to understand any consequential transformations and follow the applicable journal, institutional, disciplinary, and research-integrity requirements. ORI now highlights content provenance as increasingly important as automated and AI-enabled image manipulation becomes more capable.

09 · The Bottom Line

Authenticity Is Easier to Demonstrate When Provenance Was Preserved From the Beginning

The Bottom Line

Preserve enough original data, images, metadata, contextual records, processing history, code, and audit information to trace reported findings back to the research events that produced them. The final clean dataset or publication-ready figure is rarely sufficient by itself.

Build provenance into ordinary research practice, use secure and appropriate storage, and verify the retention rules governing your specific project. When questions arise years later, contemporaneous source records are considerably more persuasive than even the most sincere recollection of what happened.

10 · Sources and Further Reading

Authoritative Sources on Preserving Research Data and Images

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes