03 · What You Need to Know
Translate the question into temporal data requirements
Time can enter a research question in several ways. Sometimes the period is explicit: “How did student engagement change between 2020 and 2025?” Sometimes it is embedded in the design: “Does prior academic performance predict later dropout?” Sometimes the question concerns an intervention, policy, crisis, technological change, or other event and therefore requires observations before and after a meaningful date.
In each case, the first task is to translate the question into a temporal requirement before judging the dataset.
Identify what kind of time information the question requires
Consider four hypothetical questions:
- What proportion of university students currently use generative AI?
- How has generative AI use among university students changed over the past three years?
- Does generative AI use in the first semester predict academic performance in the second semester?
- Did student assessment practices change after the introduction of a new institutional AI policy?
All concern a related topic, but their temporal requirements differ substantially.
| Type of question |
Temporal requirement |
What the dataset needs |
| Current prevalence |
Data sufficiently current for the intended claim |
A relevant recent measurement period |
| Change over time |
Comparable observations across multiple periods |
Repeated measurements or comparable repeated cross-sections |
| Earlier predictor and later outcome |
Correct temporal ordering |
Predictor measured before the outcome |
| Before-and-after comparison |
Observations on both sides of the relevant event |
Measurements appropriately positioned before and after the event |
Asking merely “What years does this dataset cover?” therefore does not go far enough. You need to know what happened within those years.
Overall dataset coverage can be misleading
Suppose a longitudinal study is described as covering 2015 to 2025. That sounds ideal for a ten-year analysis.
Then you inspect the documentation and discover that the variable central to your question was introduced in 2023. The dataset spans ten years, but your variable spans three.
Different modules, instruments, assessments, biomarkers, administrative linkages, and survey questions may enter or leave a long-running study at different times. Consequently, the relevant temporal coverage is the intersection of the periods in which the variables needed for your particular analysis are available.
This is another reason to verify the specific variables required by your question rather than treating the dataset as one undifferentiated collection.
Cross-sectional data answer different temporal questions from longitudinal data
Cross-sectional data typically observe units at one point or over a relatively limited period. Longitudinal data involve repeated observations over time, although the exact design can vary.
Cross-sectional data
Provide observations representing a particular point or period and are often suited to describing conditions or associations at that time.
Longitudinal data
Follow units or otherwise provide repeated observations over time, allowing questions about trajectories, transitions, or temporal ordering that a single cross-section cannot directly provide.
A cross-sectional dataset can be excellent for estimating a prevalence or examining contemporaneous associations. It generally cannot, by itself, show within-person change over time when each participant is observed only once.
Repeated cross-sectional surveys create another possibility. They can support analyses of population-level trends when comparable samples and measures are collected at several periods, but they do not necessarily follow the same individuals. Population change and individual change are different questions.
Temporal order matters when your question implies sequence
Suppose your question asks whether an exposure predicts a later outcome. The word “predicts” may imply that the exposure precedes the outcome.
If both variables were measured simultaneously, the dataset may support an association but not the temporal ordering implied by your original formulation.
This issue becomes particularly important when researchers move casually from phrases such as “associated with” to “leads to,” “influences,” or “results in.” Temporal precedence alone does not establish causality, but a cause cannot ordinarily occur after the effect it is proposed to cause.
Watch Out
Do not infer temporal direction merely from the order in which variables appear in a model. Calling one variable a predictor and another an outcome does not establish that the predictor was measured first or occurred first.
Check dates at the variable level
For each essential variable, determine:
- when it was collected;
- which wave or waves contain it;
- what reference period the measure describes;
- whether the same participants or units were observed repeatedly;
- whether the measurement changed between waves; and
- whether the timing aligns with the other variables required by the question.
The reference period deserves particular attention. A survey conducted in January 2025 might ask about experiences “during the previous 12 months,” “during the past seven days,” or “ever.” The data-collection date and the period represented by the variable are therefore not necessarily the same thing.
This connects directly with the question of whether the variables were measured in the way your question requires. Temporal wording is part of measurement.
Before-and-after questions need meaningful observations on both sides
Suppose an institutional policy was introduced in August 2024 and you want to examine whether student behavior changed afterward. A dataset containing observations from July and September 2024 technically contains data before and after the policy.
That does not automatically make it adequate.
You would need to consider whether the observation windows are appropriate, whether other seasonal or contextual differences matter, whether measurements are comparable, and whether the design supports the inference you intend to make. A single pre-event and post-event measurement may be much weaker than a longer series when the goal is to distinguish an intervention-associated change from an existing trend.
The exact requirements depend on the design. The general principle is simpler: “before” and “after” should be defined by the research design, not merely by finding two dates that happen to fall on opposite sides of an event.
Rapidly changing phenomena make recency more consequential
Some research topics change slowly. Others can change considerably within months or a few years.
Technology adoption is an obvious example. A dataset collected before widespread adoption of a particular technology cannot directly describe contemporary use of that technology. Similarly, labor markets, public policies, educational practices, disease conditions, platform environments, and social behaviors can change enough that older observations answer a historically different question.
There is no universal age at which data become “too old.” Recency should be judged relative to the phenomenon and the claim. Historical data may be exactly what a historical question requires, while the same data may be poorly suited to estimating current prevalence.
Longitudinal coverage can shrink because of attrition
A study may begin with thousands of participants and continue for many years, but not every participant necessarily remains at every wave.
If your question requires observations from both an early and a late wave, the relevant cases are those for whom the necessary information can be linked across the required periods. Attrition can therefore reduce the usable longitudinal sample and may also affect its composition.
This links temporal coverage with the separate questions of whether enough relevant cases remain and whether missing observations create problems for the planned inference.
Changes in measurement can interrupt an apparent time series
Even when a variable appears in every wave, its meaning may not remain stable. Question wording, response categories, instruments, coding schemes, collection modes, or definitions can change.
A ten-wave dataset does not necessarily provide ten directly comparable measurements. Review questionnaires, codebooks, technical documentation, and harmonization notes before treating repeated variables as a continuous series.
If the change is substantial, you may need to restrict the analysis to a shorter comparable period, harmonize measures when methodologically defensible, or reformulate the question.
Calendar time and developmental time are not always the same
Some questions depend less on calendar year than on timing relative to an individual's or organization's trajectory.
For example, a study may ask what happens during the first year after graduation, before and after childbirth, during the first six months of employment, or following adoption of a particular technology. In such cases, the dataset needs sufficiently precise event dates or timing variables to align observations meaningfully.
A dataset covering the correct calendar years can still be unsuitable if it cannot locate observations relative to the event your question uses as its temporal anchor.
07 · A Quick Checklist
Check whether the dataset's timeline matches your question
Before using existing data for a time-dependent question, check:
Define the exact calendar period, duration, sequence, or event timing required by the research question.
Verify the actual collection dates and waves covered by the dataset.
Check which waves or periods contain every essential variable rather than relying on the dataset's overall date range.
Read the reference period in each measure, such as “past week,” “previous year,” “currently,” or “ever.”
For questions involving sequence, verify that variables occur in the chronological order implied by the research question.
For trends, confirm that measures remain sufficiently comparable across the periods being compared.
For longitudinal questions, determine whether the same units can be linked across the required observations and how many remain available.
For rapidly changing phenomena, assess whether the data are sufficiently recent for the claim you intend to make.