Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Data Were Actually Analyzed in the Research Study?

The data researchers collected are not necessarily the data that produced the published result. Trace the path from collected observations to the actual analytic dataset.

316
What Data Were Actually Analyzed? Guide 316 of 899
01 · The Question

Which observations actually produced the reported result?

A study recruited 1,000 participants, measured 40 variables, and collected data at five time points. That does not mean every participant, variable, and measurement contributed to every result.

Some participants may lack outcome data. Certain observations may be excluded according to eligibility or quality-control rules. Variables may be transformed, categorized, combined, or adjusted for. A model may use only participants with complete information on all required variables, while another analysis uses a different subset. To understand a result, you need to move beyond what data were collected and determine what data actually entered the analysis.

02 · The Short Answer

Reconstruct the analytic dataset behind the result

In Brief

The data actually analyzed are the observations, participants or other units, variables, measurements, time points, and derived values that contributed to the specific statistical or qualitative analysis producing the result.

Do not equate the recruited sample, complete database, or list of collected variables with the analytic dataset. Identify exclusions, missing data, transformations, variable definitions, repeated measurements, and analysis-specific inclusion rules because different results in the same paper may be based on different data.

03 · What You Need to Know

Trace the data from collection to analysis

Collected data and analyzed data are different things

A methods section may describe an impressive amount of data collection. Participants complete questionnaires, provide biological samples, undergo assessments, generate device measurements, or contribute administrative records. Yet the published analysis may use only part of that information.

The distinction is fundamental:

Collected data All observations and variables obtained or available to the researchers.
Analyzed data The observations and variables that actually contribute to a particular analysis after relevant processing, exclusions, missing-data handling, and variable construction.

When critically reading a result, the second category is what directly generated the estimate.

Start with the unit of analysis

First determine what constitutes an observation in the analysis. It might be a participant, patient, student, household, school, hospital, country, publication, measurement occasion, eye, lesion, classroom, or another unit.

This is not always identical to the unit that was recruited or sampled. A study may recruit 500 students nested within 20 classes, analyze repeated observations from each student, and include class-level predictors. A clinical study may enroll patients but analyze multiple lesions per patient.

Knowing the unit of analysis helps you understand the effective structure of the data and whether dependencies among observations were handled appropriately.

Find the denominator for the specific result

The headline sample size can be misleading if you treat it as the denominator for every analysis.

Suppose 1,200 participants enroll. Of these, 1,130 complete the baseline questionnaire, 990 provide the six-month outcome, and 914 have complete information for all variables in an adjusted regression model. A table reporting that regression is not based on 1,200 complete observations simply because the abstract says “N = 1,200.”

CONSORT 2025 recommends reporting who is included in each analysis and in which group. It also notes that available participant numbers may vary across outcomes and time points, making analysis-specific denominators important for interpretation.

Identify exactly which variables entered the analysis

A study may collect dozens or hundreds of variables but use only a subset in its primary model. Determine which variable represents the exposure or intervention, which outcome was actually measured, what comparator or reference category was used, and which covariates entered the model.

For observational studies, STROBE asks authors to define outcomes, exposures, predictors, potential confounders, and effect modifiers, and to report the statistical methods used to control for confounding.

Do not assume that every variable mentioned in the methods entered every model.

Raw variables may be transformed before analysis

The number recorded in the original dataset is not always the number analyzed. Researchers may log-transform skewed measurements, standardize scores, calculate change from baseline, create scale totals, derive indices, convert continuous values into categories, or combine several measurements into a composite outcome.

For example, age collected in years might enter the model as a continuous variable, age squared, five-year increments, or categories such as 18–29, 30–44, and 45 years or older. These are different analytical representations of the same original variable.

STROBE specifically asks authors to explain how quantitative variables were handled in analyses and, when variables were grouped, to report the groupings used.

Data cleaning and exclusions can change the analytic dataset

Researchers may remove duplicate records, impossible values, measurements outside prespecified quality thresholds, participants later found to be ineligible, observations collected outside permitted windows, or records failing predefined validity checks.

Some exclusions are methodologically appropriate. The critical questions are what was excluded, why, when the rule was established, and how much data were affected.

A phrase such as “data were cleaned before analysis” is therefore not enough when consequential exclusions occurred. You need to know what cleaning meant in practice.

Missing data can determine who enters a model

Missing information is one of the most common reasons the analyzed dataset differs from the recruited sample.

A complete-case analysis typically uses only observations with the required data for all variables included in that analysis. Adding one covariate with substantial missingness can therefore reduce the analytic sample even when the outcome itself is available for many more participants.

Other analyses may use multiple imputation or model-based methods that allow participants with partially observed data to contribute under specified assumptions. CONSORT 2025 emphasizes reporting both the extent of missing data and how missing data were handled because different approaches rely on different assumptions and can affect precision and bias.

Watch Out

Do not assume that a method for handling missing data somehow recreates information without assumptions. Complete-case analysis, imputation, weighting, and model-based approaches answer the missing-data problem differently and can depend on assumptions that should be considered when interpreting the result.

“N” can change from one model to the next

A paper might report:

  • 1,500 participants at baseline;
  • 1,310 with the primary outcome;
  • 1,244 in the minimally adjusted model;
  • 1,087 in the fully adjusted model;
  • 760 in a subgroup analysis.

Those are not necessarily inconsistencies. Different variables and eligibility rules can create different analytic samples. But the changes should be visible enough for you to determine what each estimate represents.

STROBE recommends reporting numbers at relevant stages of a study and indicating the number of participants with missing data for variables of interest.

Repeated measurements create another layer of data selection

Longitudinal studies may collect multiple observations from each participant. The analysis might use every available measurement, only baseline and final follow-up, measurements within a defined window, or a summary derived from several observations.

Knowing when the outcome was measured therefore does not necessarily tell you which time points entered the statistical model. Check the analytical specification.

Derived datasets may be far removed from the original records

Modern research workflows can involve substantial processing between raw data and analysis. Researchers may merge datasets, link administrative records, aggregate observations, calculate scores, identify events from codes, construct exposure histories, or apply algorithms to sensor data.

The resulting analytical variable may therefore depend on many decisions. If a study analyzes “daily physical activity,” for example, the raw accelerometer signal may have undergone rules for valid wear time, non-wear detection, intensity thresholds, aggregation, and minimum valid days before a participant receives an analytical value.

When such processing is consequential, understanding the derived variable is part of understanding the data.

Adjusted analyses use a different informational structure from crude analyses

A crude comparison may use exposure and outcome alone. An adjusted model incorporates additional variables intended to account for confounding, improve precision, or satisfy another analytical purpose.

The adjusted estimate therefore comes from a model involving more information and potentially a different subset of participants if covariates are missing.

Do not simply describe the adjusted result as though researchers repeated the same comparison with a more sophisticated statistical test. The estimand, conditioning structure, and analytic sample may differ.

The primary analysis should anchor your reconstruction

A paper may contain many models, tables, robustness checks, subgroup analyses, and alternative definitions. Begin with the primary analysis.

For that analysis, identify the participants or units, exposure or intervention, outcome, time point, comparison, covariates, transformations, and missing-data strategy. Once that dataset is clear, move to secondary analyses and ask how their data differ.

Data element Question to ask
Units What entities or observations entered the analysis?
Participants Which recruited participants contributed data?
Variables Which measured or derived variables entered the model?
Time points Which observations across time were used?
Transformations Were variables categorized, standardized, combined, or otherwise transformed?
Exclusions What observations were removed and why?
Missing data What was missing and how was it handled?
Analysis-specific N How many observations actually contributed to this result?

For qualitative research, ask what material actually entered the interpretation

The same principle extends beyond statistical analysis. Researchers may conduct interviews, collect field notes, gather documents, or record observations, but the analytic corpus may exclude particular interviews, passages, cases, or sources.

Ask what materials were transcribed, coded, categorized, compared, or otherwise used in the analysis. If some collected material was excluded or unavailable, determine why when that distinction affects interpretation.

The language changes, but the underlying question remains the same: what evidence actually produced the reported finding?

04 · A Practical Example

How 1,000 recruited participants become 742 analyzed observations

Hypothetical Example

A cohort study of sleep and academic performance

Imagine a study recruits 1,000 undergraduate students to examine whether sleep duration predicts end-of-semester academic performance.

Recruited 1,000 students consent and complete at least part of the baseline assessment.
Exposure available 930 provide sufficient wearable-device data to calculate the prespecified sleep-duration measure.
Outcome available 865 of those students have an end-of-semester academic outcome available.
Covariates available 742 have complete information for sleep duration, academic outcome, and every covariate used in the fully adjusted model.
Primary model The reported adjusted association is estimated using those 742 complete observations.

The study recruited 1,000 students, but the fully adjusted result is directly based on 742 complete observations. That distinction immediately raises useful questions: Why were data missing? Did the 742 differ from those excluded? Was complete-case analysis prespecified? Would another missing-data approach produce materially different results?

Those questions are much more informative than simply repeating “N = 1,000” from the abstract.

05 · What Researchers Often Get Wrong

Common mistakes when identifying the analyzed data

Misconception

The sample size in the abstract applies to every analysis

Different outcomes, covariates, follow-up assessments, and subgroup criteria can produce different analytic samples. Find the denominator associated with the particular result.

Misconception

Every variable collected was analyzed

Researchers often collect considerably more information than appears in the primary model. Identify the variables that actually contributed rather than treating the data-collection instrument as the analysis specification.

Misconception

Missing data only matter when the outcome is missing

Missing exposure values, covariates, baseline measures, or other model variables can also exclude participants or require analytical handling. A fully adjusted model can therefore contain fewer observations than a simpler model even when outcome availability is unchanged.

Misconception

Data cleaning is methodologically neutral

Many cleaning procedures are necessary, but exclusions, recoding rules, outlier handling, and validity thresholds can affect the analytic dataset. Consequential rules should be understood rather than treated as invisible housekeeping.

Misconception

Imputed data are equivalent to fully observed data

Imputation can be an appropriate way to address missing data, but it relies on models and assumptions. Interpretation should reflect the missing-data mechanism assumed and the sensitivity of conclusions to alternative assumptions.

06 · What This Means for You

Build a miniature data audit for the result you care about

For any important finding, you should be able to describe the data behind it in one or two sentences. State who or what entered the analysis, which measurements were used, what processing occurred, and how many observations contributed.

A simple analytic-data framework

If the reported N differs from recruitment
Trace exclusions, follow-up losses, missing variables, and other reasons for the reduction.
If N changes across models
Determine whether additional variables, subgroup restrictions, or other analysis requirements changed the analytic sample.
If variables were transformed or categorized
Record the analytical version rather than describing only the original measurement.
If missing data were imputed or modeled
Identify the method and assumptions and examine relevant sensitivity analyses.
If the paper contains many analyses
Reconstruct the dataset for the primary analysis first, then identify how secondary analytic datasets differ.

This approach prevents a common reading error: interpreting the study as though every piece of information collected from every participant flowed unchanged into the final result. In real research, the path from collection to analysis is usually more selective.

07 · A Quick Checklist

Before interpreting a result, identify the data that generated it

For the analysis you are reading, check:
Identify the unit of analysis.
Find the analysis-specific number of participants, units, or observations rather than relying on the headline sample size.
Identify the exact exposure, intervention, outcome, comparator, and covariates used.
Determine which time points or repeated measurements entered the analysis.
Identify transformations, categories, scores, composite variables, or other derived measures.
Find any data-cleaning or exclusion rules that removed observations.
Determine the extent of missing data for variables relevant to the analysis.
Identify how missing observations were handled and what assumptions the method requires.
Compare analytic sample sizes across models, outcomes, and time points rather than assuming they are identical.
08 · Frequently Asked Questions

Questions about what data were actually analyzed

What is an analytic dataset?

It is the dataset, or analysis-specific subset and representation of the data, used to perform a particular analysis. It may differ from the raw or collected dataset because of exclusions, transformations, derived variables, missing-data handling, or other analytical decisions.

Why can the sample size differ between tables in the same paper?

Different analyses may require different outcomes, covariates, time points, or eligibility conditions. Missingness can also vary by variable. Check the denominator for each analysis rather than assuming that changing sample sizes indicate an error.

What is complete-case analysis?

It generally refers to analyzing observations with the required data available for the variables involved in a particular analysis. Its validity depends on the missing-data process and analytical context, and substantial exclusions can reduce precision or introduce bias.

Does multiple imputation mean all missing data have been recovered?

No. Multiple imputation uses observed information and a statistical model to create plausible values for missing observations while representing uncertainty across imputations. Its validity depends on assumptions about the missing-data process and the specification of the imputation model.

Should I worry if adjusted and unadjusted models use different sample sizes?

You should at least determine why. Additional covariates with missing values can reduce the adjusted analytic sample, which means differences between estimates may reflect both statistical adjustment and a change in which participants contributed.

Where can I find information about the analyzed dataset?

Check participant-flow information, statistical methods, table footnotes, figure captions, supplementary material, and analysis-specific denominators. When these do not provide enough detail, identify what information is missing.

09 · The Bottom Line

The dataset collected is not necessarily the dataset analyzed

The Bottom Line

For every important result, identify the observations, variables, time points, transformations, exclusions, and missing-data procedures that actually produced the analysis rather than relying on the study's headline sample size or data-collection description.

Once you reconstruct the analytic dataset, you can see who and what the estimate truly represents. That often reveals methodological questions that remain invisible when you treat “the data” as one undifferentiated object.

10 · Sources and Further Reading

Sources and further reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes