03 · What You Need to Know
Start by defining exactly who your research question is about
Population suitability is easy to overlook because dataset descriptions often use broad labels such as “adults,” “students,” “households,” or “employees.” Your research question may use similarly broad terminology. Yet the operational definition of the population can be much narrower.
Survey methodology distinguishes the population a study intends to describe from the practical mechanisms used to identify and sample members of that population. The American Association for Public Opinion Research describes a sampling frame as the list or procedure identifying members of the survey population from which the sample is selected. Coverage error can occur when members of the target population are absent from that frame or when units outside the target population are included.
For secondary analysis, you inherit these population boundaries. You cannot redesign the original sampling frame after the data have been collected.
Translate broad population labels into explicit eligibility criteria
Suppose your proposed question asks:
How is generative AI use associated with academic engagement among university students?
Who counts as a university student for this question?
Does the intended population include undergraduate and graduate students? Public and private institutions? Full-time and part-time students? Online learners? Students in vocational or technical institutions? International students? Students across one country or several?
These are not merely semantic details. They define the population to which you intend your findings to apply.
Now compare that intended population with the dataset's actual eligibility criteria. A dataset described informally as a survey of “university students” might include only first-year undergraduates attending public universities in selected regions. That is a much more specific population.
| Population dimension |
Research question requires |
Dataset covers |
| Educational level |
All university students |
Undergraduates only |
| Year level |
All year levels |
First-year students only |
| Institution type |
Public and private |
Public universities only |
| Geographic scope |
National |
Selected regions |
The table does not automatically tell you to reject the dataset. It tells you that the question and the available population do not currently describe the same thing.
Distinguish the target population from the observed sample
Several population concepts can become relevant when evaluating an existing dataset. Terminology varies somewhat across methodological traditions, but the underlying distinctions matter.
Target population
The population about which your research question ultimately seeks to make claims.
Observed sample
The people or units whose data actually appear in the dataset and are available for your analysis.
Between those two can sit the original study's defined population, sampling frame, eligibility criteria, sampling process, recruitment procedures, and nonresponse. Each can affect how closely the observed cases correspond to the population you want to study.
Do not infer population coverage from the sample alone. Read the methodological documentation describing who was eligible, how potential participants were identified, how they were selected or recruited, and who ultimately participated.
Read the inclusion and exclusion criteria carefully
A dataset may deliberately exclude portions of a broader population. Those exclusions are not necessarily methodological flaws. They may have been entirely appropriate for the original study.
A household survey may exclude institutionalized populations. A school-based survey may omit young people who are not attending school. A health dataset may contain only people who sought care in participating facilities. An employee survey may include only organizations that agreed to participate.
The problem arises when a secondary researcher silently expands the population beyond those boundaries.
For example, evidence from currently enrolled students does not automatically describe all people of university age. Evidence from hospital patients does not automatically describe everyone with the underlying health condition. Evidence from platform users does not automatically describe people who could potentially use the platform.
Coverage problems can exclude people systematically
Population mismatch is particularly consequential when people missing from the dataset differ meaningfully from those represented.
Consider an online survey intended to describe a population in which reliable internet access is uneven. If participation effectively requires internet connectivity, some members of the intended population may have a lower probability of appearing in the data. If internet access is also related to the phenomenon being studied, the coverage problem becomes substantively important.
AAPOR identifies undercoverage as a form of coverage error in which some members of the target population are missing from the sampling frame. The important question for your study is not simply how many people are absent, but whether the mechanism producing that absence matters for the inference you want to make.
A representative sample is representative of something specific
Researchers sometimes describe a dataset as “representative” without completing the sentence. Representative of whom, under what design, and for which estimates?
A probability sample designed to represent adults living in households within a particular country may provide strong population inference for that defined population when analyzed appropriately. That does not make it representative of all adults everywhere, nor necessarily of populations explicitly excluded from the sampling frame.
When a dataset provides sampling weights, those weights may help produce estimates for the population the sampling design was intended to represent. They should not be assumed to transform the dataset into a representative sample of a different target population.
Watch Out
Do not use “nationally representative,” “population representative,” or similar language merely because the dataset is large or contains survey weights. Verify the precise population the data producer says the sampling design and weights are intended to represent.
Sample size cannot repair population mismatch
Suppose you have 100,000 observations from users of a particular online learning platform. Your question concerns all university students in the country.
The sample is enormous, but its size does not establish that platform users resemble students who do not use the platform. If platform participation is associated with institution type, socioeconomic circumstances, digital access, academic program, or other characteristics related to your outcome, the population mismatch may remain consequential regardless of N.
This is why the question of whether the dataset contains enough relevant cases should be evaluated separately from population coverage. Quantity and population fit answer different methodological questions.
Subgroup availability does not automatically support subgroup inference
A dataset may contain members of the population you care about without having been designed to represent that subgroup adequately.
Imagine a general adult survey containing 200 university students. Those observations may permit certain analyses involving students. But the mere presence of 200 students does not establish that they constitute a representative sample of university students.
Check how those cases entered the sample, whether the subgroup can be identified appropriately, whether the survey design supports subgroup estimates, and whether relevant weights or design variables apply.
Population fit also depends on the unit being studied
Not every dataset is fundamentally about individuals. Research questions may concern schools, universities, households, firms, countries, publications, courses, classrooms, hospitals, or other units.
If your question concerns universities but the dataset sampled students, you need to determine whether the design supports institution-level inference. A large number of students drawn from a small number of universities does not necessarily provide broad coverage of universities.
Likewise, thousands of patients from three hospitals are still observations arising from three hospitals. The appropriate population check should therefore begin with the unit to which the research claim applies.
Sometimes the correct solution is to narrow the question
A mismatch does not always mean abandoning the dataset. Sometimes the dataset supports a narrower question that remains substantively valuable.
If the data represent first-year students at public universities, you might redefine the target population accordingly rather than claiming to study all university students. That can be methodologically legitimate when the narrower population still addresses a meaningful research problem.
The important distinction is between transparent refinement and silently generalizing beyond the data. Existing data can reasonably influence study scope, but there are limits to how far available data should reshape the research question.
06 · What This Means for You
Make the population claim no broader than the data can support
Write down the target population before evaluating the dataset. Then compare that definition with the dataset's eligibility criteria, sampling frame, geographic coverage, recruitment or sampling procedures, exclusions, response patterns, and weighting documentation.
If the two populations differ, determine whether the mismatch is minor and manageable or whether it changes the substantive meaning of the study.
A simple decision framework
If the dataset was designed to represent the population your question concerns
Proceed while following the data producer's sampling, weighting, and variance-estimation guidance where applicable.
If the dataset represents a meaningful but narrower population
Consider narrowing the research question and conclusions to that population.
If important segments of the target population are systematically absent
Assess whether the resulting coverage problem undermines the inference you want to make.
If the dataset fundamentally represents a different population
Find another data source, redefine the question for a substantively defensible population, or reconsider the study design.
This population audit belongs alongside the broader process of checking an existing dataset before building a study around it. A dataset is not suitable merely because its topic, variables, and sample size look attractive.
07 · A Quick Checklist
Check who the dataset can actually represent
Before treating the dataset as suitable for your population, check:
Define the target population in your research question explicitly, including relevant geographic, institutional, demographic, or other boundaries.
Identify the population the original data collection was designed to cover.
Read the eligibility, inclusion, and exclusion criteria in the official study documentation.
Examine the sampling frame or recruitment mechanism and identify important groups that could be systematically omitted.
Verify the geographic, institutional, and other coverage boundaries of the dataset.
For subgroup analyses, determine whether the sampling design supports inference about the subgroup rather than merely containing some subgroup members.
If weights are provided, verify exactly which population and estimates they were designed to support.
Narrow the population claim when the dataset supports a more limited population than the original question.