01 · The Question
How many of the dataset's cases are actually relevant to your study?
A dataset contains 20,000 participants. That sounds reassuringly large. Your study, however, concerns university students aged 18 to 24 who completed a particular module, reported a particular behavior, and have valid values for the variables required by your analysis. After those conditions are applied, perhaps only 430 cases remain.
Then you divide those 430 cases into comparison groups. One group contains only 37 participants.
The number printed on the dataset's homepage was never the sample size that mattered most. For a secondary analysis, the more consequential question is whether the dataset contains enough relevant and usable cases for the particular estimates, comparisons, or models your research question requires.
03 · What You Need to Know
Count the cases your analysis can use, not merely the records in the file
Sample-size reasoning changes somewhat when you work with existing data. In a new study, you may calculate how many observations you need and then design recruitment or sampling accordingly. With an existing dataset, the available number of cases is largely fixed. Your task is to determine whether that number is adequate for the analysis you propose.
There is no universal minimum number of cases that makes a dataset suitable. Adequacy depends on the research question, estimand, expected effect or desired precision, statistical method, number and distribution of observations, subgroup sizes, missingness, and, for survey data, the sampling design.
CDC guidance similarly notes that required sample size depends on factors including the analyses to be conducted and the effect size of interest. Methodological guidance on secondary analysis recommends studying the dataset documentation closely enough to determine whether enough cases exist to generate meaningful estimates for the topic being investigated.
Distinguish the total sample from the analytic sample
The total sample is the number of observations represented in the broader dataset. Your analytic sample is the set of observations that can actually contribute to a particular analysis after the relevant criteria are applied.
Total dataset sample
All observations contained in the dataset or a particular data file.
Relevant analytic sample
Observations eligible for your research question and usable for the variables and analysis you require.
Those numbers can differ dramatically. Cases may disappear from the relevant analytic sample because they:
- fall outside the target age or population;
- belong to a wave not relevant to the question;
- were not administered the required questionnaire module;
- do not belong to the subgroup being studied;
- lack one or more variables required for the analysis;
- cannot be linked across necessary files or waves; or
- are otherwise ineligible under the study's planned inclusion criteria.
This is why the assessment should follow the earlier checks of whether the dataset contains the variables required by the question and whether those variables were measured in an appropriate way.
Subgroup questions depend on subgroup counts
Suppose a national education dataset contains 50,000 students. You want to study a particular category of students representing 1% of the sample. Before any further exclusions, that leaves approximately 500 observations.
If your analysis then compares categories within that subgroup, the relevant cells become smaller still.
A large overall sample can therefore coexist with a small sample for the question that interests you. NCES notes that the precision of sample-based estimates depends partly on the size of the subgroup for which an estimate is calculated. The headline N can be impressive while the particular estimate you need remains imprecise.
For estimation, ask how precise the result needs to be
If the purpose is to estimate a prevalence, proportion, mean, or other population quantity, “enough” is often better framed in terms of precision than a generic minimum sample size.
A smaller sample generally produces greater sampling uncertainty. For a survey estimate, this uncertainty is commonly reflected in its standard error and confidence interval. NCES explains that the margin of error depends on factors including response variation, sample size, representativeness, and subgroup size.
Thus, a dataset might technically allow you to estimate a quantity while producing a confidence interval too wide to be useful for the scientific or practical purpose of the study.
For hypothesis tests, think about statistical power and meaningful effects
For a planned inferential test, adequacy may involve statistical power: the probability of detecting an effect of a specified magnitude under defined assumptions when that effect is present.
The required sample size is not determined by the test name alone. It depends on parameters such as the effect size of interest, significance criterion, desired power, group allocation, and statistical model.
More importantly, the effect size used in such reasoning should be scientifically defensible. Choosing an unrealistically large expected effect merely because the available dataset can detect it does not make the dataset adequate for effects that would actually matter.
Watch Out
Do not declare a fixed existing sample “adequate” merely because it exceeds a popular rule of thumb. Sample-size requirements depend on the specific analysis, assumptions, desired precision, and effect size or estimand of interest.
Missing data can shrink the usable sample again
Imagine that 800 eligible cases remain after your population restrictions. If the planned model requires five variables and only 540 participants have usable values for all five, a complete-case analysis would not use 800 observations.
Nor is the problem purely numerical. The pattern and reasons for missingness can affect bias and interpretation. The separate question of whether missing data make the dataset unsuitable therefore deserves more than simply subtracting blank cells from the total.
At the feasibility stage, however, you should at least estimate how much the analytic sample could contract once the variables required by your analysis are considered together.
Complex surveys complicate the meaning of sample size
Many large surveys do not use simple random sampling. They may involve stratification, clustering, unequal probabilities of selection, and survey weights. These features affect variance estimation and therefore affect how much statistical information a nominal sample size provides for a particular estimate.
The National Center for Health Statistics describes effective sample size for complex survey estimates in relation to the design effect. For an estimated proportion, a survey design that produces greater variance than a simple random sample of the same nominal size can yield an effective sample size smaller than the observed number of respondents.
Accordingly, a complex survey with 2,000 relevant observations should not automatically be treated as statistically equivalent to 2,000 independent observations obtained through simple random sampling. Use the sampling weights, strata, clusters, replicate weights, and variance-estimation procedures specified by the data producer when they are required.
Repeated observations are not necessarily independent cases
Longitudinal and multilevel datasets introduce another distinction between rows and independent observational units. Ten measurements from each of 100 participants produce 1,000 person-time records, but not 1,000 independent people.
Similarly, students clustered within schools, patients within hospitals, or employees within organizations may share characteristics that affect the statistical information provided by the sample. The appropriate assessment depends on the planned model and data structure.
Counting rows is therefore not a substitute for understanding the unit of analysis.
Rare outcomes can leave very little information
A dataset can have many participants but few occurrences of the outcome you want to model. Suppose 10,000 participants are available, but only 45 experience the event central to your question. For some analyses, those 45 events may be more consequential than the total N of 10,000.
The same logic applies to sparse categories and cross-classifications. If a model depends on comparisons across combinations of several characteristics, inspect how observations are distributed rather than assuming that the total sample will protect every analysis from sparsity.
Enough cases does not mean the right population
Sample size and population suitability are separate questions. Ten thousand observations from the wrong population do not become suitable because the N is large.
After assessing numerical adequacy, you still need to determine whether the dataset covers the population you actually want to study. Statistical precision cannot repair a fundamental mismatch between the sampled population and the population named in the research question.
06 · What This Means for You
Estimate the usable sample before building the study around it
Before committing to an existing dataset, create a sample-feasibility map. Start with the total number of observations and progressively apply the restrictions your actual question requires. Where access permits, inspect frequencies and cross-tabulations for the variables central to the planned analysis. Methodological guidance for secondary analysis specifically recommends examining frequencies, cross-tabulations, coding patterns, and missingness before the main analysis.
Then assess whether the resulting sample is adequate for the statistical task you intend to perform. For descriptive estimation, examine expected precision. For hypothesis testing or model-based questions, use an appropriate power or sensitivity analysis where meaningful, together with the assumptions relevant to the method. For complex survey data, account for the sampling design rather than treating the observations as a simple random sample.
A simple decision framework
If the relevant analytic sample provides adequate precision or power for the intended analysis
Proceed while continuing to account for missingness, sampling design, and other methodological constraints.
If the overall sample is large but the subgroup of interest is small
Evaluate the subgroup directly rather than using the full dataset N as evidence of adequacy.
If missingness substantially reduces the usable sample
Investigate the pattern and implications of missing data before deciding whether the proposed analysis remains defensible.
If the available sample cannot support the intended precision, effect size, comparison, or model
Consider a different dataset, a defensible redesign, a narrower question, or whether the research idea should be pursued with new data.
This assessment is particularly valuable before investing substantial effort in analysis. Discovering late in the project that the subgroup central to your hypothesis contains 18 usable observations is an educational experience, but not usually the educational experience the study was designed to provide.
If the data prove too weak for the intended inference, the appropriate response may eventually be to consider whether the research idea should be abandoned or redesigned rather than progressively weakening the question until the available sample can support something.
07 · A Quick Checklist
Check the sample that your analysis will actually have
Before deciding that the dataset is large enough, check:
Identify the total number of observations and the unit of analysis.
Apply the population, age, wave, module, and other eligibility restrictions required by your research question.
Inspect how many relevant cases contain the variables required for the planned analysis.
Examine subgroup and category counts rather than relying only on the overall analytic N.
Estimate how missing data may affect the usable analytic sample.
For estimates, determine whether the expected standard errors or confidence intervals provide useful precision.
For inferential analyses, assess power or sensitivity using assumptions appropriate to the planned method and scientifically meaningful effects.
For complex survey data, review weights, strata, clusters, design effects, and the data producer's variance-estimation guidance.
Check whether repeated observations or clustering mean that the number of rows overstates the amount of independent information.