Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Can Generative AI Accurately Extract Sample Size and Participant Characteristics From a Research Paper?

AI can extract sample sizes and participant characteristics from research papers, but a study may report several different participant counts. Learn how to distinguish recruitment, enrollment, analysis, and attrition while verifying extracted information.

199
AI Extraction of Sample Size and Participants Guide 199 of 384
01 · The Question

Can AI Identify How Many Participants Were Actually Included in a Study?

You ask generative AI to extract the sample size from a research paper. It reports 350 participants. Yet the methods section says 350 individuals were recruited, the participant flow diagram shows 312 completed the study, and the primary analysis includes only 298.

Which number should appear in your literature review matrix?

Extracting participant information may appear straightforward, but research papers frequently report different numbers at different stages. The challenge is not simply finding a number. It is determining what that number represents and whether the associated participant characteristics accurately describe the population relevant to the study's findings.

02 · The Short Answer

AI Can Extract Participant Information, but It May Confuse Different Sample Counts

In Brief

Generative AI can accurately extract sample sizes and participant characteristics from research papers when the information is clearly reported and accessible. However, it may confuse recruited, enrolled, randomized, completed, and analyzed participants, misinterpret demographic tables, or overlook missing data and exclusions.

Researchers should identify which sample count is relevant to their purpose, verify its meaning against the methods and results, and record participant characteristics with their correct denominators. A single number labeled "sample size" may be insufficient when different analyses involve different participants.

03 · What You Need to Know

Why Extracting Sample Size Is More Complicated Than Finding a Number

Research papers may describe participants at several stages, from initial recruitment to the final analysis. These numbers serve different purposes, and none should automatically be substituted for another.

The appropriate count depends on what the researcher needs to understand. For instance, recruitment numbers describe the initial study population, while analytical sample sizes indicate how many observations contributed to particular findings.

Which Sample Size Should AI Extract?

Before requesting a number, distinguish the different participant counts that may appear in a paper.

Participant Count What It Represents Common Extraction Error
Screened Individuals assessed for eligibility. Reporting everyone screened as a study participant.
Eligible Individuals meeting the eligibility criteria. Assuming everyone eligible participated.
Enrolled Individuals formally included in the study according to its enrollment procedures. Confusing enrollment with final analysis.
Randomized Participants or units assigned through a random allocation procedure. Assuming every randomized participant completed the intervention.
Completed Participants who completed a specified study stage or follow-up. Assuming all completers were included in every analysis.
Analyzed Participants or observations included in a particular analysis. Reporting one analytical count as applicable to every outcome.

These categories are not necessarily applicable to every study, and terminology varies across disciplines and designs.

For example, a qualitative interview study may report the number approached, consented, interviewed, and included in the analysis. A secondary-data study may report records rather than recruited human participants.

AI should extract the relevant categories rather than force all studies into a conventional recruitment-to-analysis sequence.

Why Can a Paper Report Several Different Sample Sizes?

Participant numbers can change because of eligibility screening, nonparticipation, withdrawal, loss to follow-up, incomplete measurements, exclusions, or analysis-specific data requirements.

Consider a hypothetical survey of 500 teachers. Researchers may receive 430 responses, exclude 25 records for failing prespecified quality checks, and analyze 405 responses for descriptive statistics.

A regression model requiring complete information on additional variables might include only 376 participants.

All these numbers may be correct, but they answer different questions.

Reporting the regression sample as the total number of survey respondents would be inaccurate. Likewise, reporting the initial invitations as the number analyzed would overstate the evidence available for the statistical findings.

What Is the Difference Between Sample Size and Analytical Sample Size?

The term sample size may refer broadly to the number of participants or observations included in a study. The analytical sample size refers more specifically to those contributing to a particular analysis.

These counts can differ when data are missing or different analyses have different eligibility requirements.

For example, a study might analyze 420 responses for descriptive statistics but use 387 complete cases for multiple regression.

Some analytical approaches, including certain multiple-imputation and likelihood-based methods, can use information from participants with incomplete observations. In such cases, the number contributing information to an analysis cannot always be understood as a simple complete-case count.

AI should therefore report the analysis-specific count and the missing-data approach when relevant, rather than assuming that the smallest number mentioned in the results represents the final sample.

Can AI Extract Participant Characteristics Accurately?

Participant characteristics describe who or what was studied. Depending on the research question, they may include age, sex or gender as measured and reported, educational level, professional experience, socioeconomic characteristics, geographic setting, institutional affiliation, or other relevant attributes.

Accurate extraction requires preserving the definitions, categories, and units used by the authors.

For example, a study reporting a mean age of 21.4 years should not be summarized as involving participants aged 21 to 22 years. A mean does not describe the full age range.

Similarly, a sample consisting of 65% female participants should not automatically be described as representative of women in the target population.

Participant characteristics are descriptive information. They do not, by themselves, establish population representativeness or generalizability.

Why Do Denominators Matter When Extracting Percentages?

A percentage is meaningful only when its denominator is understood.

Suppose a paper reports that 180 participants were enrolled, but demographic information was available for only 170. If 102 of those 170 participants were female, the reported percentage is 60%.

AI might incorrectly divide 102 by 180 and report 56.7%, or repeat 60% while implying that all 180 participants provided the demographic information.

Both interpretations would misrepresent the underlying data.

Researchers should preserve the relevant numerator and denominator whenever percentages are extracted, especially when demographic variables contain missing values.

Can AI Confuse Participant Characteristics Across Study Groups?

Yes. Studies involving multiple groups often present characteristics separately for intervention and comparison groups, cohorts, sites, or subpopulations.

For example, an educational intervention might include 80 students in the intervention group and 75 in the comparison group. Their mean ages, baseline achievement scores, and prior technology experience may differ.

An AI system may incorrectly combine group-specific values or attribute characteristics from one group to the entire sample.

Researchers should check whether the extracted information describes the overall sample, a particular group, or a subgroup used in an additional analysis.

What About Clustered Studies and Different Units of Analysis?

Some studies include several levels of observation. A school-based trial might involve 12 schools, 48 teachers, and 960 students.

These are not competing sample sizes. They describe different units.

If schools were randomized but student outcomes were analyzed, the number of randomized clusters and the number of students both matter.

AI may report only the largest count and overlook the structure of the study. This can conceal information relevant to statistical independence, precision, and the interpretation of findings.

Researchers should identify the unit of recruitment, assignment, measurement, and analysis when these differ.

Can AI Extract Sample Size From Tables and Flow Diagrams?

Some systems can interpret tables and figures, but accuracy depends on the document-processing capabilities and the quality of the source.

Participant flow diagrams are particularly important because they may report exclusions, withdrawals, losses to follow-up, and analysis populations not fully described in the narrative text.

Complex PDF layouts can introduce extraction problems. Numbers may be associated with the wrong column, group, time point, or outcome.

For instance, a table might report separate sample sizes for baseline, immediate posttest, and delayed follow-up. AI could mistakenly identify the baseline count as the number completing follow-up.

Where participant counts appear in tables or diagrams, verify them visually against the original document.

What Do Reporting Guidelines Require?

Reporting guidelines provide useful references for determining which participant information should be available.

CONSORT 2025 addresses participant flow in randomized trials, including numbers assigned, receiving interventions, lost or excluded, and included in analyses. STROBE addresses reporting participant numbers, descriptive characteristics, and missing data in observational studies.

For qualitative research involving interviews and focus groups, COREQ includes reporting considerations concerning participant selection, nonparticipation, and sample description.

These guidelines are design-specific. They should not be treated as interchangeable requirements for all research papers.

They also help clarify an important distinction: incomplete reporting does not necessarily prove that a study was conducted incorrectly, but it may prevent readers from evaluating the sample adequately.

Can AI Determine Whether the Sample Size Was Adequate?

Extracting the sample size is not the same as evaluating whether it was sufficient.

A study involving 30 participants may be appropriate for one qualitative investigation but insufficient for a particular quantitative analysis. A study involving thousands of observations may still have substantial bias or inadequate representation of its target population.

Sample adequacy depends on the research question, design, analytical objectives, expected precision, relevant assumptions, and disciplinary conventions.

In quantitative research, power analysis or precision-based planning may inform sample size decisions. In qualitative research, considerations such as information power, methodological purpose, and depth of inquiry may be more relevant.

AI should not infer that a study is methodologically strong or weak from participant numbers alone. Evaluating adequacy belongs to the broader task of critically appraising the study.

What If Participant Information Is Missing or Inconsistent?

Research articles sometimes report incomplete or conflicting participant counts.

The abstract may describe 250 participants, while the methods report 248 and the results include 239. These differences may reflect exclusions or analysis-specific populations, but the explanation should be established from the paper rather than invented.

AI should identify the discrepancy and report where each number appears.

When the available text does not explain the difference, the appropriate extraction is not a guessed reconciliation. It is a transparent statement that the reported counts differ and the reason cannot be determined from the available information.

Watch Out

Do not ask AI to provide only "the sample size" when a study reports several participant counts. Specify whether you need the recruited, enrolled, randomized, completed, or analyzed population, and verify the denominator associated with every extracted characteristic.

04 · A Practical Example

Extracting Participant Information When the Numbers Change Across the Study

Hypothetical Example

A Survey of University Teachers' Generative AI Adoption

Imagine a cross-sectional study examining factors associated with university teachers' intention to use generative AI for teaching.

The researchers invite 500 teachers to participate. They receive 420 survey responses, exclude 20 responses based on stated quality criteria, and retain 400 valid responses for descriptive analysis.

Because 18 respondents have missing values on variables required for a complete-case regression analysis, that model includes 382 teachers.

Among the 400 respondents included in descriptive analysis, 240 are reported as female and 160 as male. The mean age is 39.6 years, with a standard deviation of 8.4 years.

Step 1: Separate invitations from responses

Five hundred teachers were invited, but only 420 responses were received. The invitation count should not be reported as the number of study respondents.

Step 2: Identify the valid descriptive sample

After 20 exclusions, the study retained 400 valid responses. This is the relevant sample size for the reported descriptive characteristics.

Step 3: Identify the regression sample

The complete-case regression used 382 teachers. Reporting 400 as the regression sample would overstate the number of observations included in that model.

Step 4: Verify the demographic information

The descriptive sample included 240 female respondents (60%) and 160 male respondents (40%). The mean age was 39.6 years (SD = 8.4). These characteristics describe the 400 valid responses, not necessarily the 382 complete cases used in regression.

Step 5: Record the extraction with context

"The study obtained 420 responses from 500 invited university teachers. After exclusions, 400 responses were retained for descriptive analysis, while 382 complete cases were included in regression. The descriptive sample had a mean age of 39.6 years (SD = 8.4), with 60% reported as female."

The example shows why accurate extraction requires more than selecting the most prominent participant count. Different numbers describe different stages and analytical purposes.

If the study did not report the demographic composition of the 382 complete cases, AI should not assume that it was identical to the composition of the 400-person descriptive sample.

05 · What Researchers Often Get Wrong

Common Misconceptions About AI Sample Size Extraction

Misconception

The Largest Number in the Paper Is the Sample Size

The largest count may represent individuals invited or screened rather than those enrolled or analyzed. Identify what each number represents before selecting the relevant count.

Misconception

The Sample Size Reported in the Abstract Is Always the Final Analytical Sample

Abstracts may report enrollment or overall study size while individual analyses involve fewer observations. Verify the analytical population in the methods and results.

Misconception

Every Statistical Analysis Uses the Same Participants

Missing data, subgroup restrictions, and outcome-specific eligibility can produce different analytical samples. Record the count relevant to each analysis when necessary.

Misconception

Participant Percentages Always Use the Total Study Sample

Demographic percentages may use variable-specific denominators when information is missing. Check the table notes and reported counts rather than assuming every percentage refers to the full sample.

Misconception

A Large Sample Automatically Means the Study Is Representative

Representativeness depends on sampling and participation processes, not merely sample size. A large convenience sample may still systematically exclude important segments of the target population.

Misconception

AI Can Resolve Conflicting Participant Counts by Choosing the Most Plausible Number

Conflicting counts should be investigated, not silently reconciled. If the paper does not explain the discrepancy, report the inconsistency and its source locations.

06 · What This Means for You

How to Extract Sample Sizes and Participant Characteristics Reliably With AI

The most effective approach is to treat participant information as structured data with clearly defined meanings, rather than asking AI for a single number.

This is particularly useful when preparing literature matrices, extracting evidence for systematic reviews, or comparing populations across studies.

Choose the correct participant information for your purpose

If you need to describe recruitment
Extract the numbers approached, screened, eligible, and enrolled where reported.
If you need to describe study completion
Extract the numbers completing the relevant intervention, assessment, or follow-up period.
If you need to interpret a particular result
Identify the participants or observations included in that specific analysis.
If you need demographic characteristics
Extract the reported categories, statistics, units, and denominators for the relevant sample.
If the study has multiple groups or clusters
Preserve group-specific counts and distinguish participants from higher-level units such as schools or institutions.
If the reported counts disagree
Record each count with its source location and identify any unresolved discrepancy.

A Reusable Prompt for Participant Data Extraction

Suggested Prompt

"Extract the sample size and participant characteristics from this research paper using only information explicitly supported by the document. Distinguish the numbers screened, eligible, enrolled, randomized where applicable, completed, excluded, and analyzed. Identify the sample size used for each major analysis when these differ. Extract relevant demographic and participant characteristics with their reported units, categories, and denominators. Preserve group-specific information and distinguish participants from clusters or other units. For every extracted value, identify the supporting section, table, figure, or passage. Do not calculate or infer missing values unless I explicitly request calculations. If counts conflict or information is unavailable, report the discrepancy or missing information rather than guessing."

Use a Structured Extraction Table

For research synthesis, it may be useful to record participant information in separate fields.

Extraction Field What to Record
Target population Who the researchers intended to investigate.
Sampling method How participants or units were selected.
Initial sample Recruited, enrolled, or randomized count, clearly labeled.
Final analytical sample Number included in the relevant analysis.
Attrition and exclusions Reported losses, exclusions, and reasons.
Participant characteristics Relevant demographics and other reported attributes.
Source location Section, table, figure, or page supporting the extraction.

This structure reduces ambiguity when several papers use different conventions for reporting participant numbers.

However, do not assume that every study provides all these fields. Record information as not reported when appropriate, and distinguish missing reporting from an actual absence of participants or observations.

Verify Extracted Values Against the Original Document

Check the methods section, participant flow diagram, descriptive tables, and analysis-specific notes. If AI reports a percentage, confirm its numerator and denominator. If it reports an average, confirm the associated unit and whether the value represents a mean or median.

When extracting information for a systematic review or other consequential evidence synthesis, consider independent verification by another reviewer according to the review protocol.

The Cochrane Handbook discusses data collection procedures and methods for reducing extraction errors. AI assistance does not remove the need for quality control appropriate to the review.

Accurate participant extraction also supports understanding the research design and interpreting the study's reported findings, although neither task can be completed from participant numbers alone.

07 · A Quick Checklist

Before Using AI-Extracted Sample Size and Participant Information

Verify these details against the original paper:
Identify whether each count refers to screened, eligible, enrolled, randomized, completed, or analyzed participants.
Check whether different analyses use different sample sizes.
Verify participant exclusions, attrition, and missing-data explanations.
Confirm demographic values, measurement units, and variable definitions.
Check the numerator and denominator associated with each reported percentage.
Preserve differences between intervention groups, comparison groups, and relevant subgroups.
Distinguish individual participants from clusters, institutions, or repeated observations.
Compare AI-extracted values with the methods, tables, figures, and supplementary materials.
Record unresolved inconsistencies instead of selecting a plausible number without evidence.
08 · Frequently Asked Questions

Frequently Asked Questions About AI Sample Size and Participant Extraction

Which sample size should I report in a literature review?

Report the count relevant to the information being discussed and label it clearly. For example, enrollment may describe the study population, while the analytical sample is more relevant to a particular result. When these differ substantially, reporting both may be necessary.

Can AI extract participant characteristics from demographic tables?

Yes, some systems can extract information from tables, but errors may occur when columns, groups, denominators, or units are misread. Verify all consequential values against the original table and its notes.

Is the number recruited the same as the number analyzed?

Not necessarily. Participants may withdraw, be lost to follow-up, have missing measurements, or be excluded according to the study's procedures. The recruited and analyzed counts should be distinguished when they differ.

Can AI calculate percentages from participant counts?

It can perform calculations, but the correct denominator must first be established. Any calculated percentage should be labeled as calculated rather than directly reported by the authors, and the arithmetic should be checked independently.

What if the abstract and results report different sample sizes?

Check whether the abstract describes enrollment while the results describe an analysis-specific sample. If the discrepancy remains unexplained, record both numbers and their locations rather than assuming which one is correct.

Can AI determine whether a study's sample is representative?

AI may help identify sampling procedures and characteristics relevant to representativeness, but sample size alone is insufficient. Selection mechanisms, nonresponse, coverage, and the intended target population must also be considered.

Can AI extract sample sizes from qualitative studies?

Yes. It may identify the number of participants, interviews, focus groups, cases, or other relevant units. However, these counts should not be treated as interchangeable, and sample adequacy must be evaluated in relation to the qualitative methodology.

Should I verify AI-extracted participant data manually?

Yes, particularly when the information will support a literature review, meta-analysis, or methodological comparison. Verify values against the original document and use additional independent checking when required by the review protocol.

09 · The Bottom Line

Accurate Sample Extraction Requires Understanding What Each Number Represents

The Bottom Line

Generative AI can accurately extract sample sizes and participant characteristics from research papers, but researchers must verify which participants or observations each number describes. Recruitment, enrollment, completion, and analysis counts are not automatically interchangeable.

Record participant information with its correct context, denominator, and source location. When counts differ or reporting is incomplete, preserve the distinction rather than forcing the paper into a single sample-size figure.

10 · Sources and Further Reading

Official Guidance on Participant Reporting and Research Data Extraction

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes