03 · What You Need to Know
Match the intended construct to the measurement that actually occurred
When collecting primary data, researchers can choose or design measures around their research question. Secondary analysis reverses part of that process. The measurements have already been made, often for purposes different from yours. You inherit those measurement decisions along with the data.
Methodological guidance on secondary analysis therefore emphasizes examining the original assessment tools, questionnaires, codebooks, sampling procedures, coding decisions, and other documentation. Researchers also need operational definitions for the variables they intend to analyze. The fact that a dataset contains a conveniently named column does not establish that the underlying measure corresponds to the construct required by a new study.
Start with the construct, not the variable label
Suppose you want to investigate whether frequent use of generative AI for completing university coursework is associated with academic performance.
You find a variable labelled AIUSE. Before celebrating this unusually cooperative codebook, ask what AIUSE means operationally.
Perhaps respondents were asked, “Have you ever used an artificial intelligence application?” with yes/no response options. That variable does contain information about AI use. But it does not directly measure frequency, use for coursework, use of generative AI specifically, or use during the period corresponding to the academic outcome.
Your conceptual requirement and the available measure therefore differ.
| Research question requires |
Dataset actually measures |
Potential mismatch |
| Generative AI use |
Any artificial intelligence application |
Technology is defined more broadly than the intended construct |
| Frequency of use |
Ever used: yes/no |
No information about frequency |
| Use for coursework |
Use for any purpose |
Context of use is unknown |
| Use during the relevant semester |
Ever used |
Reference periods do not align |
The variable may still support a different research question. What it cannot do is acquire information that was never collected merely because your proposed interpretation requires it.
Inspect the exact question wording and response options
For survey data, the questionnaire is often as important as the data file. SAMHSA's guidance on codebooks notes that documentation may contain exact survey questions, response codes, missing-data codes, skip patterns, frequencies, concept definitions, and information about data collection and quality.
Read the exact wording whenever possible. Small differences can materially alter interpretation:
- “Do you use generative AI?” differs from “Did you use generative AI during the past seven days?”
- “How satisfied are you with your course?” differs from a validated multi-item measure of student engagement.
- “Highest qualification” differs from years of education.
- Self-reported height differs from measured height.
- Current employment differs from employment at any point during the previous year.
Response categories matter as well. A five-category frequency measure may support analyses that a binary yes/no variable cannot. Conversely, collapsing detailed categories during analysis can sometimes be defensible, whereas recovering detail that was never collected is impossible.
Check the source of the measurement
Two variables can appear to represent the same construct while coming from different sources. Academic achievement could come from an administrative record, a standardized assessment, teacher evaluation, or participant self-report. Health status might come from clinical examination, medical records, or respondents' perceptions.
The appropriate source depends on what your research question means by the construct. Self-report is not automatically inferior, nor are administrative records automatically superior. They may simply represent different kinds of evidence and carry different forms of measurement error.
Your interpretation should remain faithful to the source. If the variable records self-reported GPA, for example, describing it simply as “GPA” can obscure a relevant measurement qualification.
Check the reference period
Time is part of measurement. “Currently,” “during the past week,” “during the previous 12 months,” and “ever” define different variables even when the substantive topic is identical.
This becomes particularly important when variables are related analytically. If a study asks whether an exposure predicts a later outcome, you need to understand when each measure was collected and what period it describes. A current outcome paired with an exposure referring to an unspecified lifetime period may not support the temporal interpretation implied by the research question.
The broader feasibility issue of whether the dataset covers the period your question requires should therefore be distinguished from the reference period built into an individual measure. A longitudinal dataset may span ten years while a particular variable still ask only about behavior during the previous seven days.
Understand how constructed variables were created
Not every variable comes directly from a single question. Datasets commonly include composite scores, indices, classifications, recoded variables, scale scores, standardized measures, and variables derived from several source items.
For a constructed variable, inspect its derivation. Ask which source items were used, how responses were combined, whether weighting or transformations were applied, how missing items were handled, and what higher or lower scores represent.
ICPSR's codebook guidance specifically recommends documenting how compiled or constructed variables were created. Without that information, a plausible-looking score may be difficult to interpret responsibly.
Do not assume a measure remains comparable across waves
Longitudinal datasets create another complication. A variable may appear in several waves without being measured identically each time.
Question wording can change. Response categories may be expanded or collapsed. Diagnostic criteria can be revised. Data-collection modes may change. A scale may be replaced or recoded.
Methodological guidance on secondary analysis specifically warns researchers to examine questionnaires and codebooks when key variables are combined across longitudinal waves because assessment and coding methods may change over time.
If your research question involves trends or change, do not assume that identical or similar variable labels establish measurement equivalence.
Watch for skip patterns and conditional questions
A variable can look as though it contains substantial missing data when some respondents were never supposed to answer the question.
Imagine a survey that first asks whether respondents have ever used generative AI. Only those answering yes receive questions about frequency. A blank frequency value among nonusers is structurally different from a user refusing to answer the frequency question.
Codebooks often document these skip patterns. Understanding them is necessary both for interpreting the measure and for later evaluating whether missing data threaten the study.
A proxy variable requires an argument, not merely a resemblance
Sometimes the exact construct you want was not measured, but a related variable is available. Using a proxy is not inherently inappropriate. The problem arises when the proxy and the intended construct quietly become synonymous in the manuscript.
Suppose household income is unavailable but eligibility for a means-tested program is recorded. That indicator may provide information related to socioeconomic circumstances, but it does not become household income. Whether it is a defensible proxy depends on the research question, institutional context, eligibility rules, theoretical rationale, and claims you intend to make.
Watch Out
Do not solve a measurement mismatch by changing the label in your manuscript. If the dataset measures perceived competence, do not report it as demonstrated competence. If it measures intention, do not call it behavior. Your terminology should follow the evidence actually recorded.
Measurement suitability is about the claim you want to make
There is rarely a meaningful answer to “Is this a good variable?” without specifying the purpose.
A binary employment indicator may be perfectly adequate for estimating the proportion of participants currently employed. It would be inadequate for a question about job quality. A single age variable may suffice for age adjustment but not for a question requiring developmental stages defined by additional criteria.
Suitability is therefore relational: this measure, for this construct, in this population, for this analysis and interpretation.
06 · What This Means for You
Audit each important measure before finalizing the question
Once you identify a candidate variable, trace it back to its measurement source. For a survey item, read the questionnaire. For a test score, inspect documentation describing the assessment and scoring. For administrative data, determine how the field was generated. For a constructed variable, examine the derivation rules.
A simple decision framework
If the available measure closely matches the construct, population, context, and reference period required
Proceed to evaluate its data quality and analytical usability.
If the measure captures the construct only approximately
Decide whether a narrower interpretation or defensible proxy is sufficient for the scientific purpose of the study.
If the measurement differs substantially from the intended construct
Revise the question or seek another source of data rather than relabelling the measure.
If the documentation is insufficient to determine what the variable means
Treat the measurement as unresolved and investigate further before building the study around it.
This evaluation should happen before the research question becomes difficult to change. Where possible, inspect the data before finalizing the research question, while preserving a clear distinction between feasibility checking and searching opportunistically for interesting results.
If the available measurement forces a substantial change in the construct, say what the study now investigates rather than retaining the more appealing original terminology. A narrower question that the data can answer is methodologically stronger than a broader question supported mainly by wishful variable naming.
07 · A Quick Checklist
Check what each important variable really measures
Before accepting a variable as suitable, check:
Write an operational definition of the construct your research question requires.
Read the exact questionnaire item, assessment description, administrative definition, or derivation rule behind the variable.
Examine response categories, units, coding decisions, score direction, and missing-value codes.
Verify who provided or generated the measurement and whether that source fits your intended interpretation.
Check the measurement's reference period and whether it aligns with the other variables in your question.
For longitudinal analyses, verify whether wording, instruments, response categories, or coding changed across waves.
Identify skip patterns or eligibility rules determining which participants were actually measured.
If using a proxy, document why it can represent the intended construct and narrow your claims accordingly.