Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Can Missing Data Make an Otherwise Attractive Dataset Unsuitable Before Analysis Even Begins?

A dataset may have the right variables, population, and time period yet lose much of its value when essential observations are missing. Learn how to assess missingness before committing to the study.

462
Can Missing Data Make a Dataset Unsuitable? Guide 462 of 603
01 · The Question

Can missing values undermine a promising dataset before you analyze anything?

You find an existing dataset that seems unusually well suited to your project. It contains the variables you need, covers the right population and period, and appears to have thousands of observations. Then you inspect the data more closely. Thirty percent of participants are missing your main outcome. A key covariate is absent for nearly half of one subgroup. In a longitudinal study, many participants who entered at baseline are no longer observed at the follow-up your question requires.

The dataset still exists. The variables still exist. The advertised sample size has not changed. Yet the evidence available for your particular research question may be much weaker than those facts initially suggest.

Missing data should therefore be considered during feasibility assessment, not discovered only after the research question, hypotheses, and analysis plan have been finalized.

02 · The Short Answer

Missing data can make an otherwise suitable dataset a poor fit

In Brief

Yes. Missing data can make an existing dataset unsuitable when too little usable information remains for the research question or when the pattern and causes of missingness create serious risks of biased or unreliable inference.

The percentage missing is only part of the assessment. You also need to determine which variables and cases are affected, whether values are genuinely missing or absent by design, how missingness overlaps across variables, and whether defensible methods are available for handling it.

03 · What You Need to Know

Missingness is a feasibility problem before it becomes an analysis problem

Missing data are common in existing datasets. Participants may decline questions, fail to complete assessments, leave longitudinal studies, become ineligible for later questions, or lack information in linked administrative records. Measurements may also be unavailable because a module was administered only to part of the sample.

The existence of missing values does not automatically invalidate a dataset. What matters is whether the remaining information, together with a defensible approach to the missingness, can support the inference required by your research question.

Research on secondary analysis emphasizes examining frequency distributions and missing-data profiles before conducting the main analysis. This preliminary inspection can reveal whether the apparent sample and variable coverage survive contact with the data themselves.

Start by asking what “missing” actually means

A blank cell can arise for very different reasons. Consider a survey that asks respondents whether they have ever used generative AI. Only those answering yes are then asked how frequently they use it.

For respondents answering no, the frequency question is intentionally skipped. Their blank frequency value does not mean that researchers unsuccessfully attempted to measure frequency. The question did not apply to them under the survey logic.

By contrast, a respondent who reported using generative AI but refused to answer the frequency question represents a different situation.

Structurally absent or not applicable The value is unavailable because the study design, questionnaire logic, eligibility rule, or sampling procedure meant that the observation was not supposed to receive that measurement.
Missing information The value would have been relevant to the observation, but usable information was not obtained or is unavailable.

Methodological guidance on secondary analysis specifically recommends distinguishing skip patterns from genuinely missing information before recoding or analyzing variables. Treating every special missing-value code as though it meant the same thing can distort both sample counts and interpretation.

Check the missingness of essential variables first

Not all missing values matter equally for your research question.

If a dataset contains 300 variables and a peripheral variable you never intend to use is 80% missing, that may have no bearing on your study. If your primary outcome is 40% missing, the implications are much more immediate.

Start with the variables identified as essential when determining whether the dataset contains what your research question requires. Examine missingness for the primary outcome, main exposure or predictor, variables defining the target population, and any variables essential to the planned model or design.

The practical question is not “Does this dataset contain missing data?” Most substantial datasets do. Ask instead, “Is the information required for my specific analysis sufficiently available?”

Do not inspect missingness one variable at a time only

Suppose five variables required by your model each have 10% missing values. It would be a mistake to assume that your analysis will therefore lose exactly 10% of the sample.

If different participants are missing different variables, a complete-case analysis requiring valid observations on all five could exclude substantially more than 10%. If the same participants are missing all five, the reduction could be much smaller.

You therefore need to examine joint missingness: which combinations of required variables are simultaneously available for the same observations.

Feasibility check Question to ask Why it matters
Variable-level missingness How much is missing from each essential variable? Identifies particularly incomplete measures
Joint missingness How many cases have the combination of variables the analysis requires? Reveals the potential analytic sample
Subgroup missingness Is missingness concentrated in particular groups? May affect comparisons and population inference
Longitudinal missingness Which participants remain observed at the required waves? Determines usable longitudinal coverage
Reason for missingness Why are the values unavailable? Influences the risk of bias and appropriate handling

Missing data can reduce the relevant sample dramatically

The advertised dataset size may bear little resemblance to the number of observations available for your model.

Imagine a dataset containing 8,000 people. Your population restrictions leave 2,100 relevant cases. Of those, 1,700 have the outcome, 1,550 have the exposure, and 1,400 have the required covariates. When all requirements are considered together, perhaps only 1,050 cases contain complete information for every variable.

Whether that is acceptable depends on the analysis. But the appropriate sample-size assessment should use the observations realistically available for the intended method, not the 8,000 records in the original dataset. This directly affects whether enough relevant cases remain for the proposed analysis.

The percentage missing does not tell you whether the missingness is harmless

Two variables can each have 20% missing data while presenting very different methodological problems.

Suppose missing income values are spread across respondents with otherwise varied characteristics. Now imagine that income is disproportionately missing among the highest-income respondents because they are particularly reluctant to report it. The numerical percentage is identical, but the implications for estimating income distributions or relationships involving income may differ considerably.

Missing-data methodology commonly distinguishes among mechanisms described as missing completely at random, missing at random, and missing not at random. These terms have technical meanings and should not be assigned casually from the observed percentage alone.

The feasibility lesson is simpler: investigate whether missingness appears related to variables, groups, outcomes, study processes, or circumstances relevant to your question. Do not treat “only 15% missing” as a methodological diagnosis.

Complete-case analysis is not automatically neutral

A common response to missing data is to analyze only observations with complete information on all required variables. This is often called complete-case or listwise analysis.

It is easy to implement, but deleting incomplete cases can reduce statistical power and may introduce bias when the complete cases differ systematically from those excluded. Research examining secondary analyses of large surveys has emphasized both risks.

This does not mean complete-case analysis is always inappropriate. Its validity depends on assumptions and context. The important point at the feasibility stage is that “we will just delete the missing cases” should not substitute for investigating how many cases would be lost and who those cases are.

Imputation does not make missingness disappear

Methods such as multiple imputation and model-based approaches can be valuable for handling incomplete data under appropriate assumptions. Their existence, however, does not mean that every incomplete dataset can be rescued.

An imputation procedure uses information and assumptions to address missing values. It does not travel backward in time and collect the measurements that were never observed. If almost everyone in a critical subgroup lacks the outcome, or an essential variable was never collected in an entire wave, the problem may be structural rather than something a routine missing-data procedure can reasonably repair.

Watch Out

Do not choose a dataset on the assumption that sophisticated imputation will solve whatever missingness you discover later. First determine why information is missing, how much usable information remains, and whether the assumptions required by the proposed method are defensible.

Missingness can be concentrated in the group you care about most

Overall percentages can conceal subgroup problems. A variable may be 8% missing across the dataset but 35% missing among the small population central to your research question.

Inspect missingness across important groups whenever feasible. This is especially relevant when your analysis compares groups or aims to estimate outcomes for populations that may have different response, participation, or record-linkage patterns.

A dataset can therefore have an apparently modest overall missing-data problem while being much less informative for the precise comparison you want to make.

Longitudinal attrition deserves special attention

In longitudinal research, missingness can accumulate across waves as participants leave the study, miss particular assessments, or re-enter later.

If your question requires baseline and five-year follow-up measurements, the initial sample size is not the relevant number. You need to know how many participants provide the required information at both points and whether those retained participants differ meaningfully from those lost to follow-up.

This connects missingness with whether the dataset truly covers the time structure your question requires. A study may technically continue for ten years while providing limited usable ten-year information for your particular variables.

Documentation is essential for interpreting special missing-value codes

Existing datasets often distinguish several forms of nonresponse or inapplicability using special codes such as “refused,” “don't know,” “not asked,” “not applicable,” “out of universe,” or other dataset-specific categories.

Do not assume that values such as 7, 8, 9, 97, 98, 99, -1, or -9 represent substantive responses or that all such codes mean generic missingness. Consult the codebook and questionnaire.

ICPSR emphasizes the importance of documentation describing variables and value labels, while methodological guidance for secondary analysis recommends reviewing coding patterns and skip logic before analysis. Poor interpretation of missing-value codes can create errors before any sophisticated statistical issue even arises.

04 · A Practical Example

When 5,000 eligible participants become 1,420 usable cases

Hypothetical Example

A promising student dataset with overlapping missingness

A researcher identifies an existing dataset with 5,000 university students and wants to examine whether academic engagement is associated with later academic performance after adjusting for prior achievement and socioeconomic background.

1. Start with 5,000 eligible students All meet the basic population criteria for the proposed question.
2. Check the outcome Later academic performance is available for 3,900 students because some participants did not complete follow-up.
3. Check the predictor Academic engagement is available for 3,600 students because the relevant questionnaire module was incomplete for some participants.
4. Add the adjustment variables Prior achievement and socioeconomic information introduce further missingness.
5. Examine joint availability Only 1,420 students have complete values for every variable required by the initially planned complete-case model.
6. Investigate who is missing Students absent from follow-up are disproportionately drawn from groups with lower baseline academic performance.
7. Reassess feasibility The researcher now evaluates the missing-data mechanism, possible analytical approaches, remaining statistical information, and whether the original question can still be answered convincingly.

The conclusion is not automatically that the dataset must be rejected. The example shows why the original sample size of 5,000 was insufficient evidence of suitability. The missing-data structure has become part of the research design.

05 · What Researchers Often Get Wrong

Common mistakes when evaluating missing data

Misconception

Every blank value represents the same kind of missing data

Values can be absent because of refusal, loss to follow-up, questionnaire routing, ineligibility, failed linkage, measurement failure, or other reasons. Those distinctions can matter for both analysis and interpretation.

Misconception

A small percentage of missing data is automatically harmless

The percentage alone does not reveal whether missingness is concentrated in an essential variable, outcome category, subgroup, or period. Its pattern and causes can be more consequential than the overall proportion.

Misconception

I can evaluate missingness separately for each variable

Individual percentages can conceal how missing values overlap. Assess how many observations contain the combination of variables required by the planned analysis.

Misconception

I can simply delete incomplete cases

Complete-case analysis may reduce sample size and can introduce bias under some missingness mechanisms. Its appropriateness depends on assumptions that should be considered rather than taken for granted.

Misconception

Multiple imputation can rescue any dataset

Imputation methods depend on information and assumptions. They do not automatically make a severely incomplete or structurally unsuitable dataset capable of answering the original question.

06 · What This Means for You

Inspect missingness while the research question can still change

When access permits, examine the data before finalizing the study. Generate frequencies for essential variables, identify special missing-value codes, examine cross-tabulations or missingness by important groups, and determine how many observations have the combinations of information your analysis requires.

This is one reason preliminary inspection before finalizing the research question can be methodologically useful. The purpose is not to search for whichever statistically interesting result happens to appear. It is to determine whether the proposed study is feasible with the evidence actually available.

A simple decision framework

If missingness is limited and the remaining information supports the intended analysis
Proceed with a prespecified and defensible strategy for handling missing data.
If missingness substantially reduces the usable sample
Reassess statistical precision, power, subgroup sizes, and the complexity of the proposed analysis.
If missingness is concentrated in important groups or appears systematically related to the study variables
Evaluate the potential for biased inference and whether an appropriate missing-data method is defensible.
If essential information is absent for most relevant cases or an indispensable group or period
Consider revising the question, finding another dataset, obtaining supplementary data, or reconsidering the study.

There is no universal percentage at which missingness automatically disqualifies a dataset. The decision depends on what is missing, why it is missing, what remains, and what your proposed inference requires.

07 · A Quick Checklist

Inspect missing data before committing to the dataset

Before finalizing the study, check:
Identify every code the documentation uses for missing, refused, unknown, not applicable, skipped, or out-of-universe values.
Calculate or inspect missingness for every variable essential to the research question.
Examine how missingness overlaps across the combination of variables required by the planned analysis.
Determine how many relevant cases remain under the analysis strategy you are considering.
Check whether missingness is concentrated in important populations, groups, outcomes, or waves.
Distinguish genuine missing information from questionnaire skips and measurements that were not applicable by design.
Investigate the plausible causes and mechanisms of missingness rather than relying on the percentage alone.
Determine whether your proposed missing-data method has assumptions that are plausible for this dataset and research question.
08 · Frequently Asked Questions

Questions about missing data in existing datasets

How much missing data is too much?

There is no universal percentage that makes a dataset acceptable or unacceptable. The implications depend on which information is missing, why it is missing, how missingness is distributed, how much usable information remains, and the assumptions of the proposed analysis.

Should I reject a dataset because it contains missing values?

No. Missing values are common in real datasets. Rejecting any dataset with missingness would eliminate many valuable sources. Evaluate whether the particular missing-data structure threatens your intended analysis.

Is “not applicable” the same as missing?

Not necessarily. A value may be absent because a question legitimately did not apply to that respondent. Consult the questionnaire logic and codebook before deciding how such observations should be represented analytically.

Can I just remove all cases with missing values?

Sometimes a complete-case analysis is defensible, but it can reduce statistical power and may introduce bias when complete cases differ systematically from excluded observations. Its appropriateness should be evaluated for the particular missing-data mechanism and analysis.

Can multiple imputation solve the problem?

Multiple imputation can be useful under appropriate conditions, but it depends on assumptions and available information. It should not be treated as a universal repair for severe, structurally missing, or poorly understood data.

Should I inspect missing data before finalizing my research question?

Yes, when access permits. Feasibility checking can reveal whether the variables and observations required by the proposed question are sufficiently available before you commit to an analysis the dataset cannot realistically support.

09 · The Bottom Line

What is missing can matter as much as what the dataset contains

The Bottom Line

Missing data can make an otherwise attractive dataset unsuitable when essential information is too incomplete, the usable analytic sample becomes inadequate, or the missingness creates serious risks for the inference your question requires.

Evaluate missingness by variable, across variables, within important groups, and across time. Distinguish genuine nonresponse from structural absence, investigate why information is missing, and decide whether a defensible analytical strategy exists before building the study around the dataset.

10 · Sources and Further Reading

Sources and further reading on missing data

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes