03 · What You Need to Know
Authenticity Depends on More Than Keeping the Final Dataset
Preserve the Original or Primary Research Data
The analytical dataset is not necessarily the primary research record.
Depending on the study, primary data may include instrument outputs, survey responses, experimental measurements, observational records, interview recordings or transcripts where appropriate, laboratory notebook entries, clinical research records, photographs, microscopy images, electrophoretic gels, sensor files, field notes, or other direct records of the research.
NIH's current intramural research guidance states that primary data should be retained, including observations and experiments that do not directly lead to publication. It specifically notes that usable confocal microscopy imaging files should generally be retained in their original format, subject to its stated qualifications for technically problematic or exceptionally large imaging data.
The general principle is powerful: do not preserve only the evidence that eventually appears in the paper.
Keep Unprocessed Image Files
For scientific images, the original unprocessed files can be essential.
ORI's forensic-image guidance states that authenticating a scientific image requires access to original data. An image discrepancy in a published figure can identify a question, but determining what happened often requires comparison with the source image.
Nature Portfolio similarly advises researchers to retain unprocessed data and metadata after publication and notes that editors may request unprocessed files during peer review or when investigating post-publication concerns.
If the only surviving version is a cropped, compressed, contrast-adjusted JPEG copied from a PowerPoint slide, much of the evidence needed to establish provenance may already be gone.
Preserve Native Formats and Relevant Metadata
Converting every file into a convenient generic format can discard information.
Native instrument or image formats may retain acquisition settings, timestamps, dimensions, channel information, calibration data, bit depth, device information, or other metadata that do not survive ordinary export.
NIH record-retention guidance for clinical research also advises keeping electronic records in native format where applicable. The specific requirement varies by research environment, but preserving native files can materially strengthen provenance.
For image research, metadata may help establish when and how an image was acquired. For computational data, metadata may define variables, units, coding schemes, versions, or provenance relationships.
Preserve Enough Context to Understand What Each File Represents
A folder containing 10,000 files named IMG001.tif through IMG10000.tif may technically preserve the images while making their scientific provenance nearly impossible to reconstruct.
Researchers need enough contextual information to determine which sample, participant, specimen, condition, time point, experiment, replicate, instrument, or procedure produced each relevant record.
This can come from structured file naming, laboratory notebooks, electronic laboratory information systems, manifests, database identifiers, metadata, sample maps, study logs, or other systems appropriate to the research.
The objective is that another qualified person should be able to connect the file to the research event it represents.
Preserve Acquisition Information When Interpretation Depends on It
For image-based research, the acquisition process can matter as much as subsequent processing.
Microscope settings, exposure times, detector parameters, objective information, acquisition software, channel settings, instrument calibration, and related information may be necessary to understand whether images are genuinely comparable.
Nature Portfolio image-integrity guidance asks authors to list image-acquisition tools and processing software and document key acquisition settings and processing manipulations in the Methods.
Not every parameter needs indefinite preservation in every project. Preserve what is necessary to interpret, reproduce, and authenticate the evidence according to the standards of the field.
Keep the Processing History Between the Original and Final Image
An original image and final figure are two ends of a chain. Researchers should be able to explain what happened between them.
Relevant records may include crop boundaries, brightness and contrast adjustments, channel combinations, pseudocoloring, nonlinear transformations, stitching, segmentation, thresholding, annotations, lane rearrangements, and figure assembly.
The amount of documentation needed depends on the operation and discipline. Routine processing may be captured in a reproducible script or standard workflow. Consequential manual edits may require more explicit documentation.
The aim is to make image processing distinguishable from unexplained image alteration.
Preserve the Full Image When Only a Crop Is Published
A crop cannot demonstrate what existed outside its boundaries.
If a paper publishes only part of a microscopy field, gel, blot, radiological image, photograph, or other image, preserve the corresponding uncropped original.
This can establish whether important information was omitted and whether the published region came from the claimed source. Nature Portfolio requires unprocessed original gel and Western blot images for accepted life-science papers and recommends retaining unprocessed data and metadata after publication.
Keep Figure-to-Source Provenance
Final figures frequently contain many panels. Months later, even their creators may struggle to remember which original file produced Panel 4C.
Maintain a mapping between figure panels and source files where practical. This might be a figure manifest, reproducible script, notebook, spreadsheet, file-naming convention, electronic lab notebook entry, or another system.
Source file
The original data or image from which the displayed evidence derives.
Figure provenance
The traceable relationship connecting the source file, processing steps, selected region or measurement, and final published panel.
This becomes particularly important when figures combine images from different experiments or files.
Preserve Raw Data Separately From Cleaned or Analytical Data
Researchers often correct errors, remove duplicates, recode variables, derive measures, and create analysis-ready datasets. Those operations need not threaten authenticity if the original record remains available and the transformations are traceable.
A useful workflow preserves source data separately and derives cleaned datasets through documented procedures. Depending on the project, this might involve scripts, database audit trails, versioned datasets, correction logs, or another reproducible process.
Overwriting the only copy of the original dataset makes later verification much harder.
Keep Data Dictionaries, Codebooks, and Variable Definitions
A numerical dataset without definitions may be impossible to interpret correctly.
Preserve information explaining variable names, units, missing-value codes, category definitions, derived variables, scoring rules, transformations, and other conventions needed to understand the data.
For qualitative research, analogous documentation may include coding frameworks, methodological memos, code definitions, and records connecting coded excerpts to source material, subject to ethical and confidentiality constraints.
Preserve Analysis and Processing Code
When results depend on software scripts, code is part of the evidentiary chain.
Preserving the exact or appropriately versioned code used to clean, transform, analyze, and visualize data can show how reported results were produced from the underlying records.
Dependencies, software versions, random seeds, configuration files, computational environments, or workflow definitions may also matter for complex analyses.
Code alone is not enough if the underlying data or documentation are missing, but it can make the path from source data to reported result substantially more transparent.
Preserve Correction and Audit Trails
Datasets legitimately change. A value may be corrected from a source document, a duplicate removed, a coding error fixed, or an identifier reconciled.
Keep enough information to show what changed and why. Depending on the system, this may include previous values, corrected values, timestamps, users, reasons for changes, source evidence, or systematic correction rules.
The principles for documenting legitimate research data corrections help distinguish a traceable correction from an unexplained alteration.
Preserve Excluded and Unpublished Data Where Required
Do not assume that only observations appearing in the final analysis matter.
Excluded observations can be important for evaluating whether exclusion criteria were applied appropriately. Additional experimental replicates can show whether a displayed image was genuinely representative. Unsuccessful or negative experiments may be relevant to reconstructing the research history.
NIH intramural guidance explicitly states that primary data from observations and experiments not directly leading to publication must also be retained under its policy.
The precise retention obligation elsewhere depends on the applicable institutional, funder, regulatory, contractual, and disciplinary requirements.
Preserve Versions of Manuscripts, Figures, and Analytical Outputs When They Matter
Research provenance extends beyond raw data.
Draft figures may show when panels were replaced or rearranged. Analysis outputs may document which model produced a reported table. Manuscript versions can establish how interpretations changed after corrections or review.
Researchers do not necessarily need to archive every autosaved draft forever. The principle is proportionality: preserve records needed to reconstruct consequential decisions and reported results.
Use Backups That Protect Against Both Loss and Silent Alteration
A file existing in one location is not a preservation strategy.
NIH intramural guidance recommends regular backup of electronic research records and archival storage designed to prevent subsequent alteration. Secure institutional storage, access controls, backup systems, checksums, immutable archives, version histories, or other mechanisms can help protect records depending on the research environment.
Storage also needs to respect participant confidentiality, security requirements, controlled-access restrictions, intellectual property, and other legal or ethical obligations. “Keep everything everywhere” is not a responsible data-management policy.
Content Provenance Is Becoming More Important as Image Generation Improves
ORI now specifically highlights content provenance as an emerging research-integrity issue in response to increasingly powerful automated and AI-enabled image manipulation. It defines content provenance as a verifiable record of where digital content originated, modifications made to it, and where it has been published.
That concept is likely to become increasingly relevant as visual evidence becomes easier to synthesize or alter convincingly.
Traditional research records remain essential, but cryptographic or machine-verifiable provenance systems may eventually supplement laboratory notebooks, metadata, audit trails, and source-file archives in some research settings.
There Is No Universal Retention Period for Every Research Record
Researchers should be cautious about memorizing one number and treating it as universal.
Retention periods vary by institution, funder, award, jurisdiction, research type, regulatory framework, intellectual-property considerations, participant protections, and whether an audit, claim, investigation, or other proceeding is ongoing.
For example, current NIH intramural guidance generally requires its intramural research records to be retained for seven years after project completion or until no longer needed for scientific reference, whichever is longer, with longer periods for some records. NIH grant recipients operate under different grant record-retention requirements, including exceptions and qualifications.
Researchers should therefore verify the policy that actually governs their project rather than importing a retention period from another institution or research context.
Watch Out
Do not wait for a journal query or integrity allegation before trying to reconstruct provenance. Once original files, metadata, processing histories, or excluded observations are lost, later explanations may be impossible to verify even when the research was entirely legitimate.