Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Can Poor Documentation Make an Existing Dataset Too Risky to Build a Study Around?

Data are difficult to reuse responsibly when you cannot determine what the variables mean, how observations were produced, or what processing occurred. Learn when documentation gaps become a feasibility problem rather than a minor inconvenience.

463
Can Poor Documentation Make a Dataset Too Risky? Guide 463 of 603
01 · The Question

What if the data are available but you cannot confidently explain what they mean?

You obtain an existing dataset that appears to contain exactly what your research question requires. The variable names look promising. The sample is large. The file opens without difficulty.

Then the questions begin. What does SCORE2 measure? Why are some observations coded -9? Was income recorded monthly or annually? Is the variable called “engagement” a questionnaire item, a composite score, or a researcher-created category? How were duplicate records handled? Which participants were eligible? What happened between the raw data and the file you received?

If the documentation cannot answer questions that materially affect interpretation or analysis, the problem is no longer administrative untidiness. It may become a threat to the defensibility and reproducibility of the study itself.

02 · The Short Answer

A dataset you cannot interpret reliably may be too risky to use

In Brief

Yes. Poor documentation can make an existing dataset unsuitable when essential information about variables, sampling, measurement, coding, processing, provenance, or file structure is missing and cannot be recovered with sufficient confidence.

Not every documentation gap is fatal. The key question is whether you can independently understand what the data represent, evaluate their suitability for your research question, reproduce the required processing, and justify your analytical decisions without relying on unsupported assumptions.

03 · What You Need to Know

Documentation is part of the evidence you need to interpret the data

A data file does not explain itself. A column of numbers becomes scientifically meaningful only when you know what was measured or recorded, from whom or what, under which conditions, using which definitions, at what time, and with what transformations or coding decisions.

This is why established data repositories place substantial emphasis on documentation. ICPSR's current data-submission requirements state that deposits should contain the data files, documentation, and study-level information needed for others to independently understand, evaluate, and reuse the data. Its examples include codebooks, data dictionaries, questionnaires, user guides, variable and value labels, descriptions of derived variables, and syntax needed to reproduce, merge, weight, or interpret data.

The FAIR principles make a related point from the perspective of data reuse: data and metadata should be richly described with information that allows potential users to determine whether a dataset is useful in a particular context, including its scope, provenance, processing, limitations, and relevant variable definitions.

Start with the questions the documentation must let you answer

The amount of documentation required depends on the dataset and the proposed study. A simple table may require less explanation than a multiwave survey assembled from dozens of files. The standard should therefore not be “Does this dataset have a codebook?” but “Can the available documentation answer the questions necessary for my intended use?”

For many quantitative secondary analyses, you may need to establish:

  • who or what the observations represent;
  • how participants or units were identified, sampled, or recruited;
  • when the information was collected;
  • what each essential variable means;
  • how values and missing values are coded;
  • how important variables were measured;
  • how derived variables or scores were constructed;
  • how files and repeated observations can be linked;
  • whether weights, strata, clusters, or other design variables are required;
  • what cleaning, transformations, exclusions, or processing produced the released data; and
  • which version of the dataset you are actually using.

If your analysis depends on any of these details, missing documentation can prevent you from evaluating whether the dataset is suitable in the first place.

A codebook and a data dictionary are useful, but they may not be enough

A well-prepared codebook or data dictionary can explain variable names, labels, response categories, units, special codes, and other variable-level information. But some questions require study-level or procedural documentation.

For example, a variable label may tell you that WEIGHT is a survey weight. It may not tell you which population it targets, whether different weights are required for different modules, or how variance should be estimated. A variable called TOTAL_SCORE may have an informative label but still require documentation explaining which items contributed to the score and how missing items were handled.

Documentation What it may help establish Why you may need it
Codebook or data dictionary Variable names, labels, values, units, missing codes Interpret and recode variables correctly
Questionnaire or instrument Exact wording, response options, routing, scale items Evaluate what was actually measured
Methodology or technical report Population, sampling, recruitment, collection procedures Evaluate population coverage and study design
Processing documentation or syntax Cleaning, derivations, merges, transformations Understand how released values were produced
Weighting documentation Weight construction and intended application Analyze complex survey data appropriately
Version information Release, corrections, changes between files Identify exactly which data support the study

Variable labels can hide consequential ambiguity

Suppose the dataset contains a variable labelled “academic performance.” That does not tell you whether it represents cumulative GPA, current-semester GPA, standardized assessment performance, self-reported grades, or a composite created by the original investigators.

If no documentation resolves the ambiguity, you cannot confidently determine whether the variable was measured in the way your question requires.

Guessing from the distribution is not a satisfactory substitute. A variable ranging from 1 to 5 might be a Likert-type item, a five-category classification, an averaged scale, or something else entirely. Plausibility is not provenance.

Undocumented codes can quietly corrupt an analysis

Consider a variable with values from 1 to 5 plus 9. If the documentation is missing, you might treat 9 as a legitimate high score. In the original study, however, 9 may have meant “don't know,” “not applicable,” or “missing.”

Likewise, a value of zero could mean none, below a detection threshold, not measured, structurally inapplicable, or an actual numerical zero depending on the dataset.

These problems are especially consequential when evaluating whether missing data threaten the proposed study. You cannot characterize missingness accurately if you do not know which codes represent missing information.

Derived variables require a traceable construction process

Existing datasets frequently contain scales, indices, standardized scores, risk classifications, imputed values, aggregated measures, or variables generated from several source fields.

For an important derived variable, you should ideally be able to determine how it was produced. ICPSR specifically identifies descriptions of computed or derived variables and relevant syntax or code among the documentation that can support independent understanding and reuse.

Without such information, interpretation can become speculative. A variable called “SES_INDEX” may sound straightforward, but which indicators were included? Were they standardized? Weighted? How were missing components treated? Was a higher score intended to represent greater or lower socioeconomic advantage?

If that index is peripheral, uncertainty may be manageable. If it is the central exposure or outcome, the same documentation gap may be much more serious.

Provenance matters when data have passed through several hands

Provenance describes where data came from and what happened to them along the way. This becomes particularly important when you receive a dataset from a colleague, laboratory, organization, institutional office, previous student, or project team rather than directly from the original repository or data producer.

The FAIR principles emphasize detailed provenance for reusable data, including information about who generated or collected the data and how they were processed or transformed.

Imagine receiving a file called survey_final_FINAL_clean2.xlsx. It contains fewer rows than the sample described in the original report and several variables that do not appear in the questionnaire. Without a record of exclusions, recoding, merges, or derived variables, you may not know how the file relates to the original data.

The filename may be an archaeological artifact. Your methods section should not have to become one.

Version information can be methodologically important

Datasets can be corrected, updated, expanded, harmonized, or re-released. Variables may change between versions, errors may be fixed, and documentation may be revised.

Record the version, release date, persistent identifier, or other identifying information supplied by the data producer whenever available. FAIR guidance specifically recommends documenting the version of archived or reused data.

This helps another researcher determine which exact data supported your results and protects you from accidentally combining documentation from one release with data from another.

Poor documentation can prevent you from checking other feasibility criteria

Documentation problems rarely remain isolated. If eligibility rules are unclear, you may not know whether the dataset covers your target population. If wave dates are unclear, you may not know whether the required period is actually covered. If variable construction is unclear, you cannot confidently evaluate measurement suitability.

Documentation is therefore not simply another item on a feasibility checklist. It is often the evidence required to complete the rest of that checklist.

Not every documentation gap makes a dataset unusable

A missing README file is not automatically fatal. Nor does a dataset need documentation in one prescribed format.

The information may be distributed across a codebook, questionnaire, methodology report, repository record, publication, syntax file, data dictionary, and correspondence with the data producer. ICPSR, for example, makes study documentation available separately from data access for many collections, allowing researchers to evaluate variables and study characteristics before obtaining restricted data.

The issue is whether consequential uncertainty can be resolved from authoritative information.

Manageable documentation gap Missing information is peripheral or can be resolved reliably from authoritative documentation, syntax, repository records, or the data producer.
Consequential documentation gap Uncertainty affects the meaning of essential variables, sample definition, measurement, processing, missingness, weighting, linkage, or another feature necessary to justify the study.

Contacting the data producer can resolve uncertainty, but document the answer

If an essential detail is missing, contacting the repository or data producer can be entirely appropriate. ICPSR itself directs users to seek clarification when information about variables or released files is unclear.

If the clarification materially affects your analysis, preserve it in a form that can inform your methods and reproducibility records. An undocumented verbal assurance may help you personally understand the file but provides little traceability for someone attempting to evaluate or reproduce your study later.

Watch Out

The more consequential an undocumented assumption is to your findings, the less comfortable you should be building the study around it. “This probably means...” is a weak foundation for an essential variable, population definition, weighting decision, or derived outcome.

04 · A Practical Example

When a promising dataset becomes difficult to defend

Hypothetical Example

A student dataset with an undocumented engagement score

A researcher receives a dataset from a previous institutional project containing 6,000 student records. The proposed study examines whether student engagement predicts academic performance.

1. Identify the apparent variables The file contains ENG_SCORE, GPA, demographic variables, and several institutional characteristics.
2. Search the available documentation A spreadsheet defines GPA and the demographic fields but says only “engagement score” for ENG_SCORE.
3. Look for the measurement source The original questionnaire cannot be located, and no documentation identifies the items contributing to the engagement score.
4. Investigate processing No syntax or methodological note explains whether ENG_SCORE is a sum, mean, weighted index, standardized score, or researcher-created classification.
5. Seek authoritative clarification The original data manager can explain several administrative variables but cannot reconstruct how the engagement score was created.
6. Make the feasibility decision Because engagement is the study's central predictor and its meaning cannot be established confidently, the researcher decides not to build the proposed study around that variable.

The dataset may remain useful for questions involving well-documented variables. The documentation problem does not necessarily invalidate every possible use of the file. It makes this particular study difficult to defend because the uncertainty concerns the construct at its center.

05 · What Researchers Often Get Wrong

Common mistakes when working with poorly documented data

Misconception

I can infer what a variable means from its name

Names and labels can be abbreviated, outdated, or ambiguous. An essential variable should be interpreted using authoritative documentation about its definition, measurement, coding, or construction.

Misconception

The values look reasonable, so the coding must be obvious

Apparently reasonable values can contain special missing codes, categories, transformed values, or units that are not apparent from the data alone. Plausibility checks are useful, but they do not replace documentation.

Misconception

If the dataset has already been used in research, I do not need to inspect its documentation

Previous use can provide useful context but does not establish suitability for your question or guarantee that you have the same file, version, variables, population, or processing decisions as previous researchers.

Misconception

Documentation affects reproducibility but not the validity of my own analysis

Poor documentation can prevent you from understanding measurements, missing values, sampling, weights, exclusions, or derived variables correctly. Those are immediate validity concerns, not merely future reproducibility concerns.

Misconception

I should reject any dataset with incomplete documentation

Documentation quality exists on a continuum. The relevant question is whether missing information creates consequential uncertainty for your intended analysis and whether that uncertainty can be resolved reliably.

06 · What This Means for You

Treat unresolved documentation gaps as research risks

Before committing to an existing dataset, create a short documentation audit alongside your variable and sample checks. For every feature essential to the study, identify the authoritative source that tells you what it means or how it was produced.

If you cannot identify that source, flag the uncertainty rather than filling the gap from intuition.

A simple decision framework

If essential variables and study procedures are clearly documented
Proceed to verify that the documented information matches what appears in the data.
If a documentation gap concerns a peripheral feature
Determine whether the study can proceed without relying on that information.
If important information is absent but recoverable from an authoritative source
Resolve and document the uncertainty before finalizing the research design.
If unresolved ambiguity affects an essential variable, sample definition, processing decision, or analytical requirement
Consider another dataset, revise the study to avoid the undocumented element, or reconsider whether the evidence is strong enough for the proposed question.

When access permits, inspect the data before finalizing the research question and compare what you observe with what the documentation says should be present. Documentation and data should be read together. Neither should be treated as a substitute for the other.

If major uncertainties remain unresolved, the appropriate response may eventually be to ask whether the available evidence is simply too weak to support the proposed study convincingly. An attractive dataset is not automatically worth the methodological debt created by assumptions you cannot verify.

07 · A Quick Checklist

Check whether the documentation is good enough to support reuse

Before building a study around existing data, check:
Verify that every essential variable has an authoritative definition, label, coding scheme, unit, or measurement description.
Locate the questionnaire, instrument, administrative definition, or other source needed to understand how central variables were measured.
Confirm that special values, missing-data codes, skips, and not-applicable categories are documented.
Identify how important derived variables, scores, classifications, or indices were constructed.
Review documentation for the population, sampling or recruitment process, eligibility criteria, and data-collection period.
Determine how files, waves, repeated observations, or external records should be linked when your analysis requires them.
Verify the purpose and correct application of weights, strata, clusters, or other design variables when relevant.
Record the dataset version, release information, identifier, and source so the exact data used can be traced.
Resolve consequential ambiguities with the repository or data producer rather than filling them with assumptions.
08 · Frequently Asked Questions

Questions about poorly documented existing data

Does every dataset need a formal codebook?

Not necessarily in one particular format, but you need sufficient authoritative documentation to understand and evaluate the information required by your study. The relevant details may be distributed across a data dictionary, questionnaire, methodology report, user guide, syntax files, and repository metadata.

Can I infer undocumented variable meanings from previous publications?

Previous publications can provide useful clues and context, but they may not document the exact variable, file version, coding, or processing used in your copy of the data. Prefer documentation from the data producer or repository when available.

What should I do if a variable has no label?

Search the codebook, questionnaire, data dictionary, syntax, technical reports, and other authoritative study materials. If an essential variable remains uninterpretable, do not assign it a meaning based solely on its values or position in the file.

Can I contact the original researcher or data producer for clarification?

Yes. Clarification from the data producer or repository can resolve important uncertainties. When the information affects your methods or interpretation, preserve enough documentation to make the decision traceable.

Is poor documentation itself evidence that the data are low quality?

Not necessarily. The underlying observations may have been collected carefully even when documentation is incomplete. The practical problem is that inadequate documentation can prevent a secondary researcher from evaluating, interpreting, and reusing those observations confidently.

When is poor documentation serious enough to reject a dataset?

Consider rejecting or redesigning around the dataset when unresolved documentation gaps affect information essential to your research question, such as the meaning or construction of a central variable, the population represented, important processing decisions, missing-data codes, linkage procedures, or the correct analytical design.

Why does the dataset version matter?

Datasets may be corrected or revised over time. Recording the exact version or release helps ensure that your documentation matches the data analyzed and allows others to identify the evidence underlying your results.

09 · The Bottom Line

If you cannot explain the data, you may not be ready to build a study around them

The Bottom Line

Poor documentation can make an existing dataset too risky when unresolved uncertainty prevents you from understanding essential variables, evaluating the study population and measurements, reconstructing important processing decisions, or justifying the analysis.

Look beyond the presence of a codebook and ask whether the available documentation lets you independently understand and evaluate the evidence your question depends on. Minor gaps may be manageable; consequential ambiguity should be resolved before the study proceeds rather than converted into an undocumented assumption.

10 · Sources and Further Reading

Sources and further reading on dataset documentation

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes