01 · The Question
Does having data on your topic mean you have the data your question needs?
You find an existing dataset that seems almost made for your project. It covers the general topic, comes from a credible source, includes thousands of observations, and perhaps has already supported published research. Then you begin planning the analysis and discover that one variable central to your research question is not actually there.
This is an easy problem to overlook because researchers often evaluate datasets at the level of topics rather than variables. A survey described as covering education, health, employment, or technology may contain extensive information about that domain without containing the specific pieces of information your proposed analysis requires.
Before building a study around existing data, therefore, you need to ask a more precise question: Can every essential concept in my research question be represented using variables that are actually available to me?
03 · What You Need to Know
Translate the research question into specific data requirements
The most useful way to evaluate variable availability is to work backward from the research question. Instead of opening a codebook and asking, “What interesting variables are here?”, first ask, “What information would have to be present for me to answer my question?” This keeps the proposed analysis question-driven rather than allowing an attractive dataset to quietly redefine the study.
Guidance from ICPSR similarly recommends identifying data requirements before evaluating possible datasets and then assessing whether those requirements are met. Research on secondary analysis also emphasizes developing an analytic plan that specifies the variables to be considered and examining the available documentation carefully.
Break the question into its essential concepts
Suppose your proposed question is:
Is students' use of generative AI for academic work associated with their academic performance after accounting for prior achievement and socioeconomic background?
That sentence already implies several data requirements. You would need some representation of generative AI use, academic performance, prior achievement, and socioeconomic background. Depending on the intended analysis, you might also need identifiers, demographic characteristics, institutional variables, sampling variables, weights, or other covariates.
Do not stop at identifying the apparent independent and dependent variables. Ask what the proposed design and analysis require. A question involving change needs information that permits change to be examined. A subgroup comparison requires a variable identifying the relevant groups. An adjusted association requires the proposed adjustment variables. A multilevel question may require variables identifying the relevant clusters.
| Part of the proposed question |
Information needed |
What to verify in the dataset |
| Generative AI use |
Exposure or predictor |
A variable that represents the intended form of AI use |
| Academic performance |
Outcome |
A usable measure of the intended performance outcome |
| Prior achievement |
Adjustment variable |
A baseline or earlier measure appropriate to the proposed adjustment |
| Socioeconomic background |
Adjustment variable |
One or more variables that can reasonably represent the intended construct |
This exercise produces a variable requirements list. That list becomes the bridge between the conceptual research question and the actual dataset.
Separate essential variables from desirable variables
Not every variable you would like to have is equally important. Some are indispensable because the research question cannot be answered without them. Others would strengthen interpretation or permit additional analyses but are not necessary for the central question.
A useful distinction is between essential variables and desirable variables. If the dataset lacks a desirable variable, the study may remain viable with an acknowledged limitation. If it lacks an essential variable, the proposed question may no longer be answerable as written.
Essential variable
Without it, the central research question or required analysis cannot be carried out as intended.
Desirable variable
Its presence would improve the study, permit additional adjustment, or enrich interpretation, but its absence does not necessarily invalidate the central analysis.
This distinction matters because otherwise a researcher can become trapped between two unhelpful extremes: abandoning a workable dataset because it is not perfect, or accepting a fundamentally inadequate dataset because it contains many interesting variables.
Check the codebook, not merely the dataset description
A dataset's landing page or catalogue description usually operates at the study level. It may tell you the population, broad subjects, data-collection period, methodology, and access conditions. Those details are useful, but they do not establish that the specific variables you need are present.
Variable-level documentation is more informative. ICPSR describes a codebook as documentation of a data collection's contents, structure, and layout, including details for individual variables. The UK Data Service similarly recommends documentation that identifies variable names, descriptions, values, and other information needed to understand secondary data.
Look for resources such as:
- codebooks;
- data dictionaries;
- variable lists;
- questionnaires or interview schedules;
- user guides;
- file descriptions;
- technical documentation; and
- documentation for restricted-use files, when relevant.
Some major data providers make this process relatively transparent. For example, the U.S. National Health and Nutrition Examination Survey provides searchable variable lists alongside questionnaires, datasets, documentation, and codebooks. The UK Data Service catalogue also supports variable searching and directs users to study documentation for question wording and data dictionaries.
Verify that the variable is in the version you can actually use
Finding a variable in documentation does not always mean it will appear in the file available to you. Large studies may distribute several files, waves, modules, or access versions. Some variables may exist only in restricted-use data because detailed geography, sensitive information, or other potentially identifying information has been withheld from public files.
ICPSR specifically recommends comparing public- and restricted-use codebooks when the variables required by a research question may not be present in the public file. A variable may therefore exist somewhere within the broader data collection but still be unavailable for your planned project.
This is where variable suitability intersects with whether access depends on permission from another party. For feasibility purposes, the relevant question is not merely whether a variable exists somewhere. It is whether you can legitimately obtain and use it under the conditions of your study.
Watch Out
Do not design the study on the assumption that a restricted, linked, unreleased, or separately held variable will eventually become available. If an essential variable depends on access you have not secured, treat that access as a feasibility condition rather than an established fact.
A variable name is not enough
Suppose the codebook contains a variable labelled “AI use.” That sounds promising, but what does it actually record? Any AI use? Frequency of use? Use for academic work? Use of generative AI specifically? Teacher-directed use? Self-reported use during the previous week?
The variable's presence answers only the first feasibility question. You must still determine whether the variable was measured in the way your question requires. Two datasets may both contain a variable called “income,” “achievement,” “stress,” or “technology use” while operationalizing those concepts quite differently.
For this reason, inspect the original question wording, response categories, units, coding scheme, reference period, skip patterns, and derivation rules whenever these affect your interpretation. The distinction is subtle but consequential: variable availability asks whether the information exists; measurement suitability asks whether that information represents the construct in the way your question requires.
Check whether the variables can be used together
A second complication arises when all the required variables appear somewhere in the data collection but cannot straightforwardly be analyzed together.
For example, one variable may appear only in one survey wave while another appears only in a later wave. Variables may be stored in separate files that require a valid linking identifier. One measure may apply only to adults while another was administered only to adolescents. A module may have been given to a subsample rather than the full study population.
Thus, “the dataset contains both variables” is not always sufficient. You need to establish that the required variables coexist for the observations relevant to your proposed analysis.
This is also why variable checking cannot completely replace the broader task of evaluating an existing dataset before building a study around it. Variable availability is one feasibility criterion among several.
Presence does not guarantee usable information
A variable can technically exist yet contribute very little usable information. It might contain extensive missing values, apply only to a small subsample, have almost no variation, or be suppressed for many observations.
Those are related but distinct feasibility problems. Once you establish that the necessary variables exist, examine whether enough relevant cases are available and whether missing data threaten the planned analysis.
The sequence matters. There is little reason to calculate an elaborate power analysis for a proposed model if one of the model's essential variables does not exist in the dataset at all. Secondary-data projects have their own version of “measure twice, cut once.” The codebook is usually cheaper than the regret.
06 · What This Means for You
Make variable feasibility an explicit decision before committing to the study
Before finalizing a project based on existing data, create a simple variable map. Place each concept required by the research question on one side and the candidate dataset variable on the other. Record where the variable appears, what documentation defines it, whether it is accessible, and whether any uncertainty remains.
If you cannot identify a credible candidate variable for an essential concept, do not bury that problem in the methods section later. Resolve it while the study is still being designed.
A simple decision framework
If every essential concept has an accessible candidate variable
Proceed to evaluate how those variables were measured and whether they are analytically usable.
If a desirable but nonessential variable is absent
Consider whether the study remains defensible without it and document the resulting limitation.
If an essential variable exists only in restricted or permission-controlled data
Treat access as unresolved until the relevant conditions and approvals are understood.
If an essential concept has no defensible variable
Revise the question, find another dataset, collect additional data, or reconsider the project rather than pretending a weak substitute answers the original question.
When you have access to the data, it can also be useful to inspect the dataset before finalizing the research question. Documentation tells you what variables are supposed to be present; preliminary inspection can reveal whether their actual contents create additional feasibility concerns.
There is nothing inherently wrong with refining a question in response to the data that genuinely exist. Secondary analysis necessarily involves some negotiation between substantive interests and available information. The methodological problem begins when the available variables determine the question so completely that the researcher starts searching for any relationship that can be reported. Keeping an explicit variable requirements list helps preserve the distinction between reasonable adaptation and letting available data drive question shopping.
07 · A Quick Checklist
Check variable availability before you commit
Before building the study around an existing dataset, check:
Break the research question into every concept and piece of information the proposed analysis requires.
Distinguish variables that are essential to the central question from variables that would merely be desirable.
Search the official codebook, data dictionary, variable list, questionnaire, and user documentation rather than relying on the dataset description.
Record the exact variable names or identifiers corresponding to each required concept.
Verify that each required variable is included in the specific file, wave, module, and access version you can use.
Confirm that variables needed together are available for compatible cases and can be linked when stored in separate files.
Flag any proposed proxy variable and justify why it can represent the intended concept before treating it as an acceptable substitute.
If an essential variable is absent, decide explicitly whether to revise the question, seek another dataset, obtain additional data, or reconsider the project.