03 · What You Need to Know
Documentation is part of the evidence you need to interpret the data
A data file does not explain itself. A column of numbers becomes scientifically meaningful only when you know what was measured or recorded, from whom or what, under which conditions, using which definitions, at what time, and with what transformations or coding decisions.
This is why established data repositories place substantial emphasis on documentation. ICPSR's current data-submission requirements state that deposits should contain the data files, documentation, and study-level information needed for others to independently understand, evaluate, and reuse the data. Its examples include codebooks, data dictionaries, questionnaires, user guides, variable and value labels, descriptions of derived variables, and syntax needed to reproduce, merge, weight, or interpret data.
The FAIR principles make a related point from the perspective of data reuse: data and metadata should be richly described with information that allows potential users to determine whether a dataset is useful in a particular context, including its scope, provenance, processing, limitations, and relevant variable definitions.
Start with the questions the documentation must let you answer
The amount of documentation required depends on the dataset and the proposed study. A simple table may require less explanation than a multiwave survey assembled from dozens of files. The standard should therefore not be “Does this dataset have a codebook?” but “Can the available documentation answer the questions necessary for my intended use?”
For many quantitative secondary analyses, you may need to establish:
- who or what the observations represent;
- how participants or units were identified, sampled, or recruited;
- when the information was collected;
- what each essential variable means;
- how values and missing values are coded;
- how important variables were measured;
- how derived variables or scores were constructed;
- how files and repeated observations can be linked;
- whether weights, strata, clusters, or other design variables are required;
- what cleaning, transformations, exclusions, or processing produced the released data; and
- which version of the dataset you are actually using.
If your analysis depends on any of these details, missing documentation can prevent you from evaluating whether the dataset is suitable in the first place.
A codebook and a data dictionary are useful, but they may not be enough
A well-prepared codebook or data dictionary can explain variable names, labels, response categories, units, special codes, and other variable-level information. But some questions require study-level or procedural documentation.
For example, a variable label may tell you that WEIGHT is a survey weight. It may not tell you which population it targets, whether different weights are required for different modules, or how variance should be estimated. A variable called TOTAL_SCORE may have an informative label but still require documentation explaining which items contributed to the score and how missing items were handled.
| Documentation |
What it may help establish |
Why you may need it |
| Codebook or data dictionary |
Variable names, labels, values, units, missing codes |
Interpret and recode variables correctly |
| Questionnaire or instrument |
Exact wording, response options, routing, scale items |
Evaluate what was actually measured |
| Methodology or technical report |
Population, sampling, recruitment, collection procedures |
Evaluate population coverage and study design |
| Processing documentation or syntax |
Cleaning, derivations, merges, transformations |
Understand how released values were produced |
| Weighting documentation |
Weight construction and intended application |
Analyze complex survey data appropriately |
| Version information |
Release, corrections, changes between files |
Identify exactly which data support the study |
Variable labels can hide consequential ambiguity
Suppose the dataset contains a variable labelled “academic performance.” That does not tell you whether it represents cumulative GPA, current-semester GPA, standardized assessment performance, self-reported grades, or a composite created by the original investigators.
If no documentation resolves the ambiguity, you cannot confidently determine whether the variable was measured in the way your question requires.
Guessing from the distribution is not a satisfactory substitute. A variable ranging from 1 to 5 might be a Likert-type item, a five-category classification, an averaged scale, or something else entirely. Plausibility is not provenance.
Undocumented codes can quietly corrupt an analysis
Consider a variable with values from 1 to 5 plus 9. If the documentation is missing, you might treat 9 as a legitimate high score. In the original study, however, 9 may have meant “don't know,” “not applicable,” or “missing.”
Likewise, a value of zero could mean none, below a detection threshold, not measured, structurally inapplicable, or an actual numerical zero depending on the dataset.
These problems are especially consequential when evaluating whether missing data threaten the proposed study. You cannot characterize missingness accurately if you do not know which codes represent missing information.
Derived variables require a traceable construction process
Existing datasets frequently contain scales, indices, standardized scores, risk classifications, imputed values, aggregated measures, or variables generated from several source fields.
For an important derived variable, you should ideally be able to determine how it was produced. ICPSR specifically identifies descriptions of computed or derived variables and relevant syntax or code among the documentation that can support independent understanding and reuse.
Without such information, interpretation can become speculative. A variable called “SES_INDEX” may sound straightforward, but which indicators were included? Were they standardized? Weighted? How were missing components treated? Was a higher score intended to represent greater or lower socioeconomic advantage?
If that index is peripheral, uncertainty may be manageable. If it is the central exposure or outcome, the same documentation gap may be much more serious.
Provenance matters when data have passed through several hands
Provenance describes where data came from and what happened to them along the way. This becomes particularly important when you receive a dataset from a colleague, laboratory, organization, institutional office, previous student, or project team rather than directly from the original repository or data producer.
The FAIR principles emphasize detailed provenance for reusable data, including information about who generated or collected the data and how they were processed or transformed.
Imagine receiving a file called survey_final_FINAL_clean2.xlsx. It contains fewer rows than the sample described in the original report and several variables that do not appear in the questionnaire. Without a record of exclusions, recoding, merges, or derived variables, you may not know how the file relates to the original data.
The filename may be an archaeological artifact. Your methods section should not have to become one.
Version information can be methodologically important
Datasets can be corrected, updated, expanded, harmonized, or re-released. Variables may change between versions, errors may be fixed, and documentation may be revised.
Record the version, release date, persistent identifier, or other identifying information supplied by the data producer whenever available. FAIR guidance specifically recommends documenting the version of archived or reused data.
This helps another researcher determine which exact data supported your results and protects you from accidentally combining documentation from one release with data from another.
Poor documentation can prevent you from checking other feasibility criteria
Documentation problems rarely remain isolated. If eligibility rules are unclear, you may not know whether the dataset covers your target population. If wave dates are unclear, you may not know whether the required period is actually covered. If variable construction is unclear, you cannot confidently evaluate measurement suitability.
Documentation is therefore not simply another item on a feasibility checklist. It is often the evidence required to complete the rest of that checklist.
Not every documentation gap makes a dataset unusable
A missing README file is not automatically fatal. Nor does a dataset need documentation in one prescribed format.
The information may be distributed across a codebook, questionnaire, methodology report, repository record, publication, syntax file, data dictionary, and correspondence with the data producer. ICPSR, for example, makes study documentation available separately from data access for many collections, allowing researchers to evaluate variables and study characteristics before obtaining restricted data.
The issue is whether consequential uncertainty can be resolved from authoritative information.
Manageable documentation gap
Missing information is peripheral or can be resolved reliably from authoritative documentation, syntax, repository records, or the data producer.
Consequential documentation gap
Uncertainty affects the meaning of essential variables, sample definition, measurement, processing, missingness, weighting, linkage, or another feature necessary to justify the study.
Contacting the data producer can resolve uncertainty, but document the answer
If an essential detail is missing, contacting the repository or data producer can be entirely appropriate. ICPSR itself directs users to seek clarification when information about variables or released files is unclear.
If the clarification materially affects your analysis, preserve it in a form that can inform your methods and reproducibility records. An undocumented verbal assurance may help you personally understand the file but provides little traceability for someone attempting to evaluate or reproduce your study later.
Watch Out
The more consequential an undocumented assumption is to your findings, the less comfortable you should be building the study around it. “This probably means...” is a weak foundation for an essential variable, population definition, weighting decision, or derived outcome.
07 · A Quick Checklist
Check whether the documentation is good enough to support reuse
Before building a study around existing data, check:
Verify that every essential variable has an authoritative definition, label, coding scheme, unit, or measurement description.
Locate the questionnaire, instrument, administrative definition, or other source needed to understand how central variables were measured.
Confirm that special values, missing-data codes, skips, and not-applicable categories are documented.
Identify how important derived variables, scores, classifications, or indices were constructed.
Review documentation for the population, sampling or recruitment process, eligibility criteria, and data-collection period.
Determine how files, waves, repeated observations, or external records should be linked when your analysis requires them.
Verify the purpose and correct application of weights, strata, clusters, or other design variables when relevant.
Record the dataset version, release information, identifier, and source so the exact data used can be traced.
Resolve consequential ambiguities with the repository or data producer rather than filling them with assumptions.