03 · What You Need to Know
Trace the data from collection to analysis
Collected data and analyzed data are different things
A methods section may describe an impressive amount of data collection. Participants complete questionnaires, provide biological samples, undergo assessments, generate device measurements, or contribute administrative records. Yet the published analysis may use only part of that information.
The distinction is fundamental:
Collected data
All observations and variables obtained or available to the researchers.
Analyzed data
The observations and variables that actually contribute to a particular analysis after relevant processing, exclusions, missing-data handling, and variable construction.
When critically reading a result, the second category is what directly generated the estimate.
Start with the unit of analysis
First determine what constitutes an observation in the analysis. It might be a participant, patient, student, household, school, hospital, country, publication, measurement occasion, eye, lesion, classroom, or another unit.
This is not always identical to the unit that was recruited or sampled. A study may recruit 500 students nested within 20 classes, analyze repeated observations from each student, and include class-level predictors. A clinical study may enroll patients but analyze multiple lesions per patient.
Knowing the unit of analysis helps you understand the effective structure of the data and whether dependencies among observations were handled appropriately.
Find the denominator for the specific result
The headline sample size can be misleading if you treat it as the denominator for every analysis.
Suppose 1,200 participants enroll. Of these, 1,130 complete the baseline questionnaire, 990 provide the six-month outcome, and 914 have complete information for all variables in an adjusted regression model. A table reporting that regression is not based on 1,200 complete observations simply because the abstract says “N = 1,200.”
CONSORT 2025 recommends reporting who is included in each analysis and in which group. It also notes that available participant numbers may vary across outcomes and time points, making analysis-specific denominators important for interpretation.
Identify exactly which variables entered the analysis
A study may collect dozens or hundreds of variables but use only a subset in its primary model. Determine which variable represents the exposure or intervention, which outcome was actually measured, what comparator or reference category was used, and which covariates entered the model.
For observational studies, STROBE asks authors to define outcomes, exposures, predictors, potential confounders, and effect modifiers, and to report the statistical methods used to control for confounding.
Do not assume that every variable mentioned in the methods entered every model.
Raw variables may be transformed before analysis
The number recorded in the original dataset is not always the number analyzed. Researchers may log-transform skewed measurements, standardize scores, calculate change from baseline, create scale totals, derive indices, convert continuous values into categories, or combine several measurements into a composite outcome.
For example, age collected in years might enter the model as a continuous variable, age squared, five-year increments, or categories such as 18–29, 30–44, and 45 years or older. These are different analytical representations of the same original variable.
STROBE specifically asks authors to explain how quantitative variables were handled in analyses and, when variables were grouped, to report the groupings used.
Data cleaning and exclusions can change the analytic dataset
Researchers may remove duplicate records, impossible values, measurements outside prespecified quality thresholds, participants later found to be ineligible, observations collected outside permitted windows, or records failing predefined validity checks.
Some exclusions are methodologically appropriate. The critical questions are what was excluded, why, when the rule was established, and how much data were affected.
A phrase such as “data were cleaned before analysis” is therefore not enough when consequential exclusions occurred. You need to know what cleaning meant in practice.
Missing data can determine who enters a model
Missing information is one of the most common reasons the analyzed dataset differs from the recruited sample.
A complete-case analysis typically uses only observations with the required data for all variables included in that analysis. Adding one covariate with substantial missingness can therefore reduce the analytic sample even when the outcome itself is available for many more participants.
Other analyses may use multiple imputation or model-based methods that allow participants with partially observed data to contribute under specified assumptions. CONSORT 2025 emphasizes reporting both the extent of missing data and how missing data were handled because different approaches rely on different assumptions and can affect precision and bias.
Watch Out
Do not assume that a method for handling missing data somehow recreates information without assumptions. Complete-case analysis, imputation, weighting, and model-based approaches answer the missing-data problem differently and can depend on assumptions that should be considered when interpreting the result.
“N” can change from one model to the next
A paper might report:
- 1,500 participants at baseline;
- 1,310 with the primary outcome;
- 1,244 in the minimally adjusted model;
- 1,087 in the fully adjusted model;
- 760 in a subgroup analysis.
Those are not necessarily inconsistencies. Different variables and eligibility rules can create different analytic samples. But the changes should be visible enough for you to determine what each estimate represents.
STROBE recommends reporting numbers at relevant stages of a study and indicating the number of participants with missing data for variables of interest.
Repeated measurements create another layer of data selection
Longitudinal studies may collect multiple observations from each participant. The analysis might use every available measurement, only baseline and final follow-up, measurements within a defined window, or a summary derived from several observations.
Knowing when the outcome was measured therefore does not necessarily tell you which time points entered the statistical model. Check the analytical specification.
Derived datasets may be far removed from the original records
Modern research workflows can involve substantial processing between raw data and analysis. Researchers may merge datasets, link administrative records, aggregate observations, calculate scores, identify events from codes, construct exposure histories, or apply algorithms to sensor data.
The resulting analytical variable may therefore depend on many decisions. If a study analyzes “daily physical activity,” for example, the raw accelerometer signal may have undergone rules for valid wear time, non-wear detection, intensity thresholds, aggregation, and minimum valid days before a participant receives an analytical value.
When such processing is consequential, understanding the derived variable is part of understanding the data.
Adjusted analyses use a different informational structure from crude analyses
A crude comparison may use exposure and outcome alone. An adjusted model incorporates additional variables intended to account for confounding, improve precision, or satisfy another analytical purpose.
The adjusted estimate therefore comes from a model involving more information and potentially a different subset of participants if covariates are missing.
Do not simply describe the adjusted result as though researchers repeated the same comparison with a more sophisticated statistical test. The estimand, conditioning structure, and analytic sample may differ.
The primary analysis should anchor your reconstruction
A paper may contain many models, tables, robustness checks, subgroup analyses, and alternative definitions. Begin with the primary analysis.
For that analysis, identify the participants or units, exposure or intervention, outcome, time point, comparison, covariates, transformations, and missing-data strategy. Once that dataset is clear, move to secondary analyses and ask how their data differ.
| Data element |
Question to ask |
| Units |
What entities or observations entered the analysis? |
| Participants |
Which recruited participants contributed data? |
| Variables |
Which measured or derived variables entered the model? |
| Time points |
Which observations across time were used? |
| Transformations |
Were variables categorized, standardized, combined, or otherwise transformed? |
| Exclusions |
What observations were removed and why? |
| Missing data |
What was missing and how was it handled? |
| Analysis-specific N |
How many observations actually contributed to this result? |
For qualitative research, ask what material actually entered the interpretation
The same principle extends beyond statistical analysis. Researchers may conduct interviews, collect field notes, gather documents, or record observations, but the analytic corpus may exclude particular interviews, passages, cases, or sources.
Ask what materials were transcribed, coded, categorized, compared, or otherwise used in the analysis. If some collected material was excluded or unavailable, determine why when that distinction affects interpretation.
The language changes, but the underlying question remains the same: what evidence actually produced the reported finding?