01 · The Question
How much should you look at existing data before committing to the question?
With newly collected data, the ideal sequence can seem straightforward: formulate the research question, design the study, collect the data, and analyze them. Existing data complicate that sequence. The dataset already exists, and some of its characteristics may need to be inspected before you can know whether your proposed study is feasible.
You might need to determine whether essential variables are populated, whether enough relevant cases exist, whether categories are extremely sparse, whether variables can be linked across files, or whether missingness leaves enough usable observations. Some of these questions cannot be answered confidently from a catalogue description or codebook alone.
But looking at the data also creates another possibility: you see an interesting association and then formulate a question or hypothesis around it. That can be legitimate exploratory research when reported transparently, but it is different from testing a question specified independently of the observed pattern.
The issue is therefore not simply whether you should inspect the data. It is what you inspect, why you inspect it, and what claims you later make about the resulting analysis.
03 · What You Need to Know
Separate feasibility inspection from substantive result inspection
Existing-data research creates a genuine design tension. You need enough knowledge about the dataset to avoid committing to an impossible study, yet excessive knowledge of the relationships you plan to test can influence which hypotheses, variables, exclusions, and analyses you choose.
Methodological work on preregistration of preexisting data recognizes this problem explicitly. Researchers using secondary data cannot change many features of the original study, and prior knowledge of the data may influence hypotheses and analytical decisions. Recommended transparency practices therefore include reporting whether the dataset was previously accessed or analyzed and what relevant information was already known.
Some inspection is simply feasibility checking
Imagine proposing a study involving three variables in an existing dataset. The documentation confirms that all three variables exist, but it does not tell you whether they contain substantial missingness or whether enough observations have valid information on all three simultaneously.
Opening the dataset to inspect frequencies and missingness may be necessary to decide whether the study can proceed. That is different from estimating the association among the three variables and then deciding whether the relationship looks interesting enough to study.
Feasibility inspection
Examines whether the data can support the proposed study, such as variable availability, coding, sample counts, missingness, file structure, ranges, or linkage.
Substantive result inspection
Examines the relationships, effects, group differences, model results, or other patterns that the eventual research question or hypothesis would evaluate.
The boundary is not always perfectly sharp. A frequency distribution may itself reveal substantively interesting information. A scatterplot created to identify impossible values may also reveal a relationship between variables. What matters is being explicit about what was inspected and recognizing when that inspection has supplied knowledge relevant to the eventual hypothesis.
Inspect variable availability when documentation cannot settle it
Documentation should normally be your first source for determining whether the dataset contains the variables required by your question. But the data file may reveal additional feasibility issues.
A variable documented in the codebook may contain values for only one survey module. A supposedly available measure may be almost entirely empty in the population you intend to analyze. Variables stored in separate files may require identifiers that do not link as expected.
Checking these features before finalizing the question can prevent substantial wasted work.
Inspect counts rather than assuming the advertised sample is your sample
If the dataset contains 30,000 records but your question concerns a narrow subgroup, you may need to count how many observations actually satisfy the relevant eligibility criteria.
Then check how many of those cases contain the variables required by the proposed analysis. This provides a more defensible basis for determining whether enough relevant cases are available.
This type of inspection does not require testing the hypothesis. You can often evaluate eligibility counts, subgroup frequencies, and joint variable availability without estimating the substantive relationship at the center of the proposed study.
Inspect missingness before committing to a model
A codebook can tell you that a variable exists. It cannot always tell you how much usable information it contains for your particular analytic sample.
Preliminary inspection may therefore include missing-value frequencies, patterns of joint availability across required variables, questionnaire skip codes, and retention across longitudinal waves. These checks can reveal whether missing data threaten the feasibility of the study.
The purpose at this stage is not necessarily to choose whichever missing-data strategy produces the preferred result. It is to understand the structure of the evidence before designing an analysis around it.
Inspect coding and distributions for impossible or unusable values
Basic inspection can also reveal discrepancies between documentation and the data file. You may find undocumented codes, impossible values, unexpected category labels, duplicated identifiers, or variables with virtually no variation.
A proposed predictor that has the same value for almost everyone may technically exist while offering little information for the intended analysis. Similarly, an outcome with an extremely rare category may change what analyses are realistic.
These are legitimate feasibility concerns. They should be distinguished from repeatedly trying alternative outcomes, predictors, transformations, or subgroup definitions until an attractive relationship emerges.
Prior access does not automatically invalidate secondary analysis
Researchers sometimes assume that once they have opened a dataset, preregistration or confirmatory analysis becomes meaningless. The methodological literature takes a more nuanced position.
Preregistration of secondary analyses can still restrict later analytical flexibility and make decisions more transparent. However, researchers should report relevant prior access and prior knowledge of the dataset. The credibility of a confirmatory interpretation depends partly on what was already known when the hypothesis and analysis plan were specified.
A useful record can include:
- when you obtained access to the dataset;
- which files or variables you inspected;
- which descriptive summaries you examined;
- whether relationships among variables were inspected;
- whether you or collaborators previously analyzed the same data; and
- what relevant results were already publicly available.
This kind of transparency allows readers to judge how independent the eventual hypothesis was from the observed data.
Preregistration is particularly useful when the data already exist
Preregistration is sometimes misunderstood as something that must occur before data exist. For secondary analysis, that is impossible by definition.
Instead, researchers can specify the question, hypotheses, operationalizations, exclusions, preprocessing, and planned analyses before conducting the focal analyses. Templates developed specifically for secondary data also ask researchers to report how the data were obtained, whether they were previously explored, and what relevant information is already known.
This does not erase prior exposure. It makes the sequence of decisions more visible and can reduce the opportunity to alter the analysis after seeing the focal results.
Watch Out
Preregistering after examining the relationship you intend to test does not make that prior inspection disappear. Record relevant prior knowledge honestly. Transparency about a data-informed hypothesis is methodologically stronger than presenting it as though it was specified without knowledge of the result.
Exploratory analysis is not the problem
Exploration is a legitimate part of research. Existing datasets are often valuable precisely because they allow researchers to discover patterns and generate new questions.
The methodological problem arises when exploratory and confirmatory reasoning are blurred. If inspection of the data generates a hypothesis, that hypothesis can be reported as data-informed or exploratory and investigated further, ideally with independent evidence when strong confirmatory claims are desired.
Methodological discussions of preexisting data emphasize distinguishing hypothesis-generating exploration from hypothesis testing because using the same observed pattern both to generate and to test a hypothesis changes how conventional inferential results should be interpreted.
You can sometimes protect the focal analysis from unnecessary exposure
If the dataset is large enough, one option is to conduct feasibility checks or develop analysis code using a subset of the data while reserving other observations for the focal analysis. Another approach is to validate the analytical pipeline on simulated data. In controlled-access settings, data custodians may also provide metadata, descriptive information, or restricted preliminary access before releasing the full analytical file.
These approaches do not suit every dataset or research design, and splitting data reduces the information available for each stage. Their value depends on the scientific goal. They can nevertheless help when researchers want to resolve practical analytical issues without repeatedly examining the focal results.
Write down the purpose of inspection before you inspect
A simple discipline can help: specify in advance what you need to learn from the data.
For example:
- verify that the target subgroup contains at least enough observations for the planned feasibility assessment;
- determine the proportion of missing values in essential variables;
- confirm that participant identifiers link the required files;
- check whether an ordinal variable contains the categories described in the codebook; or
- verify which survey waves contain usable observations.
Once those questions are answered, stop. If you subsequently decide to explore substantive relationships, that can be done explicitly as exploration rather than smuggled into what was supposed to be a feasibility check.
07 · A Quick Checklist
Inspect existing data without losing track of what you have learned from them
Before inspecting data to finalize a research question, check:
Write down the proposed research question or substantive area before examining the analytical data.
Use codebooks, questionnaires, technical reports, and other documentation before opening the data when they can answer the feasibility question.
Specify exactly what you need to inspect, such as counts, coding, missingness, ranges, linkage, or wave availability.
Avoid inspecting focal associations, group differences, model results, or outcome patterns when they are unnecessary for feasibility checking.
Keep a record of which files, variables, summaries, and analyses you or your collaborators have already seen.
Report relevant prior access or knowledge when preregistering or describing the eventual analysis.
If inspection generates the question or hypothesis, label that process transparently rather than reconstructing it as an a priori prediction.
Consider preregistering the focal analysis after feasibility is established and before examining the relationships it is intended to test.