Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Can Generative AI Invent or Misreport Statistical Results?

Generative AI can invent numerical results, extract the wrong values, or misinterpret statistics from real studies. Researchers should trace every consequential AI-generated statistic to the original analysis, dataset, table, figure, or publication.

38
AI Hallucinations of Statistical Results Guide 38 of 80
01 · The Question

Can AI Give You a Statistical Result That Was Never Actually Reported?

You ask a generative AI system to summarize a study's findings. It tells you that the treatment significantly improved the outcome, perhaps adding a p-value, effect size, confidence interval, or percentage difference.

The numbers look perfectly ordinary. Nothing about p =.032 or a 95% confidence interval of 1.14 to 2.67 announces that something is wrong.

But can generative AI invent such values?

Yes. It can generate numerical results that were never reported, extract the wrong number from a real source, alter a value, confuse one statistic with another, or correctly reproduce a statistic while misinterpreting what it means. For researchers, statistical plausibility is not enough. The number needs provenance.

02 · The Short Answer

Can Generative AI Hallucinate Statistics?

In Brief

Yes. Generative AI can invent or misreport statistical results, including sample sizes, percentages, means, coefficients, p-values, confidence intervals, effect sizes, test statistics, and other quantitative findings.

It can also accurately reproduce a number while attaching the wrong interpretation to it. Before using an AI-generated statistic in research, trace the value and its meaning to the original analysis output, dataset, table, figure, supplementary material, or publication.

03 · What You Need to Know

How Can AI Get Statistical Results Wrong?

A Plausible Number Is Easy to Generate

Statistical reporting follows recognizable patterns. Research articles routinely contain means and standard deviations, percentages, coefficients, test statistics, p-values, confidence intervals, effect sizes, sample sizes, and other quantitative information.

A language model can learn those patterns. It can therefore generate a statement that looks statistically conventional without independently establishing that the number came from an actual analysis.

This is the numerical version of a broader problem with AI hallucination in research. A response can have the correct form while lacking the factual basis that the form appears to imply.

Statistically plausible The value looks reasonable for that kind of analysis or result.
Empirically verified The value can be traced to the actual data, analysis, or reported study result.

The first does not establish the second.

AI Can Invent a Result Entirely

The clearest failure occurs when the model provides a number for which there is no underlying analysis or source.

Suppose you ask, “What was the effect size in this study?” If the paper does not report one, an appropriate response might be that no effect size was found in the supplied information, perhaps followed by an explanation of whether one could be calculated from available data.

A hallucinated response might instead provide a value such as Cohen's d = 0.62 and describe it as a moderate effect.

The interpretation of 0.62 might be mathematically conventional. The problem is more basic: where did 0.62 come from?

Unless the value was reported or transparently calculated from verified inputs, it is not a study result.

AI Can Extract the Wrong Number From the Right Source

Numerical errors do not always involve fabrication. Scientific papers contain many numbers, often densely packed across text, tables, figures, appendices, and supplementary files.

A model may locate a genuine number but assign it to the wrong variable, group, time point, or outcome. It might report the baseline sample size as the final analytic sample, confuse a standard deviation with a standard error, or take the control group's mean as the intervention group's mean.

Recent evidence on LLM-assisted data extraction illustrates why quantitative information deserves special scrutiny. A 2026 systematic review of 27 studies evaluating LLMs for evidence-synthesis data extraction reported substantial heterogeneity in performance, with numerical data generally extracted less reliably than categorical and string variables. The review concluded that current evidence supports assistive workflows with human verification rather than autonomous extraction.

Likewise, Gougherty and Clipp found very high accuracy for several discrete and categorical ecological data-extraction tasks but substantially poorer performance for certain quantitative data, leading them to emphasize additional quality assurance.

The practical lesson is not that every AI-extracted number is suspect. It is that source access does not make numerical extraction infallible.

AI Can Change a Number Slightly and Still Change the Conclusion

Some numerical errors are visually tiny but inferentially consequential.

Consider a reported p-value of.053. If an AI summary changes it to.035, only two digits have moved. Under a conventional.05 significance threshold, however, the generated account may now describe a result as statistically significant when the reported analysis did not cross that threshold.

The same problem can occur with confidence intervals. Changing an interval from 0.94–1.21 to 1.04–1.21 may alter whether a null value is contained in the interval, depending on the effect measure.

Small transcription errors are therefore not necessarily small scientific errors.

AI Can Report a Correct Number With the Wrong Label

A numerical value is inseparable from what it represents.

A value of 0.45 could be a correlation coefficient, regression coefficient, standardized effect, probability, proportion, factor loading, or something else entirely. A number copied correctly but assigned to the wrong statistic becomes misinformation.

What may be correct What AI may get wrong Possible consequence
0.42 Calling a correlation an effect size from an intervention The study appears to demonstrate an effect it did not estimate
n = 312 Calling the recruited sample the final analytic sample Attrition or exclusions disappear
p =.041 Attaching it to the wrong predictor The wrong relationship appears statistically significant
95% CI [0.70, 0.91] Assigning it to the wrong effect estimate Uncertainty is represented incorrectly
68% Describing a subgroup percentage as the overall prevalence The magnitude of the finding is distorted

AI Can Confuse Reported Statistics With Calculated Statistics

Sometimes a requested statistic is not directly reported but can legitimately be calculated from information in the paper. This creates an important distinction.

If an AI system transparently calculates a value using verified inputs and an appropriate formula, the resulting number is a derived calculation. It should be described as such.

If the system simply supplies a number and presents it as though the original authors reported it, the provenance has been misrepresented.

Reported result A statistic explicitly presented by the original researchers.
Derived result A statistic calculated later from reported or otherwise verified data.

Those categories should not be silently collapsed. In meta-analysis and evidence synthesis, for example, calculated effect sizes may be entirely appropriate, but the calculation method and source values need to be traceable.

Correct Statistics Can Still Receive Incorrect Interpretations

Not every statistical hallucination requires a false number. The interpretation can be wrong even when the value is copied perfectly.

An AI system might describe a statistically significant result as “important” or “large” without evidence about practical magnitude. It might treat a nonsignificant p-value as proof that there is no effect. It may describe a confidence interval as the range containing 95% of individual observations or interpret correlation as causal evidence.

Those are statistical interpretation errors rather than fabricated values, but the research consequence can be similar: the generated prose tells the reader something the analysis does not establish.

This is why AI-generated scientific claims require scrutiny even when the numbers themselves are genuine.

AI Can Generate Internally Inconsistent Statistics

Numbers in a research report are often mathematically related. A reported test statistic, p-value, confidence interval, sample size, and effect estimate may constrain one another.

A generated response can violate those relationships. It may provide a confidence interval inconsistent with the point estimate, a percentage inconsistent with the stated numerator and denominator, or a p-value that does not plausibly correspond to the reported test statistic and degrees of freedom.

Arithmetic consistency checks can catch some such errors. They cannot establish that a result is real, because a fabricated set of numbers could also be internally consistent.

Watch Out

Mathematical consistency is useful for detecting errors, but it is not provenance. A perfectly coherent set of invented statistics remains invented.

Exact Numbers Deserve More Verification, Not Less

Specificity often makes generated information feel authoritative. A statement that “the intervention improved outcomes” may invite skepticism. Add β = 0.31, p =.004, and it suddenly looks as though someone must have run an analysis.

That impression is precisely why numerical claims deserve direct verification.

The more exact the generated statistic, the easier the verification question becomes: where, exactly, does this number come from?

04 · A Practical Example

How One Incorrect p-Value Can Change the Story of a Study

Hypothetical Example

A result becomes “significant” during summarization

Suppose you ask an AI system to summarize the findings of a paper comparing two instructional approaches.

Original paper For the primary outcome, the paper reports a group difference with p =.058 and discusses the evidence cautiously.
AI-generated summary “The intervention produced a statistically significant improvement compared with the control condition (p =.038).”
What went wrong The generated p-value does not match the paper, and the changed value leads the AI to characterize the finding differently.
Why it matters If you use the generated summary in a literature review, the study now appears to provide evidence that its reported analysis did not provide under the stated threshold.
Researcher action Check the original results table and surrounding interpretation, record p =.058, and represent the finding according to the study and the inferential framework being used rather than the AI summary.

The lesson extends beyond p-values. A wrong denominator changes a percentage. A wrong sign reverses a coefficient. A misplaced decimal can transform an effect size. Statistical information should be traced, not merely trusted because it looks statistically literate.

05 · What Researchers Often Get Wrong

Common Mistakes When Using AI-Generated Statistics

Misconception

A Precise Number Is Less Likely to Be Hallucinated

Precision is a feature of the output, not evidence of provenance. Generative AI can produce decimal places, confidence limits, sample sizes, and p-values whether or not those values came from a real analysis.

Misconception

If AI Has the Paper, It Will Copy the Numbers Correctly

Providing the source can improve accuracy, but numerical extraction can still fail through omissions, transcription errors, field confusion, or selection of the wrong value. Empirical evaluations of LLM-assisted scientific extraction show that performance varies substantially by task and data type.

Misconception

If the Arithmetic Works, the Result Must Be Genuine

Internal consistency can identify some mistakes but cannot prove that a statistic came from the claimed source. Invented numbers can be mathematically coherent.

Misconception

A Correct p-Value Means the Interpretation Is Correct

A p-value does not establish effect magnitude, practical importance, causal inference, or the probability that a hypothesis is true. AI can reproduce a value accurately while describing its meaning incorrectly.

Misconception

AI Can Safely Fill in a Statistic the Authors Forgot to Report

AI may calculate a missing statistic when the necessary verified inputs and an appropriate method are available, but that calculation must be transparent. It should not invent a value or present a derived result as though the original researchers reported it.

Misconception

Numerical Errors Are Minor Compared With Fabricated Papers

A single incorrect value can reverse the apparent direction, significance, magnitude, or uncertainty of a finding. Statistical hallucinations can therefore alter the substantive interpretation of real research.

06 · What This Means for You

Treat Every Important AI-Generated Number as Traceable Data

Researchers do not need to manually reanalyse every statistic mentioned by AI. They do need to establish where consequential values came from.

The verification target depends on the task. If AI is summarizing a published study, compare the statistic with the relevant table, figure, results paragraph, appendix, or supplementary file. If it is analyzing your own dataset, compare the generated narrative with the actual statistical output and analysis code. If it calculated a new value, verify both the inputs and calculation.

A simple decision framework

If AI reports a statistic from a paper
Locate the exact value in the original source and confirm its variable, group, time point, and analysis.
If AI reports a statistic from your own analysis
Compare it directly with the software output, code, or verified dataset rather than the generated narrative.
If AI calculates a statistic that was not reported
Verify the input values, formula, assumptions, and calculation, and label the value as derived rather than reported.
If AI interprets a statistic
Evaluate the interpretation separately from whether the number itself was copied correctly.
If you cannot trace the number
Do not present it as an empirical result.

A useful rule is that every number important enough to appear in your research should have a route back to its origin. “The AI gave it to me” is not that route.

07 · A Quick Checklist

Before Using an AI-Generated Statistical Result

For every consequential statistic, check:
Locate the exact value in the original results, analysis output, dataset, table, figure, or supplementary material.
Confirm what statistic the number represents and which variable, group, model, or time point it belongs to.
Verify sample sizes and denominators rather than checking percentages alone.
Check signs, decimal places, confidence limits, units, and other details that can materially change interpretation.
Distinguish statistics reported by the original researchers from statistics calculated later.
Evaluate the interpretation separately from the accuracy of the numerical value.
Check related statistics for obvious internal inconsistencies, while remembering that consistency alone does not establish authenticity.
Do not publish or rely on a numerical result whose provenance you cannot establish.
08 · Frequently Asked Questions

Frequently Asked Questions About AI-Generated Statistics

Can AI invent a p-value?

Yes. A language model can generate a plausible-looking p-value even when no corresponding statistical test was performed. Verify the value against the actual analysis or source.

Can AI invent an effect size or confidence interval?

Yes. It can generate either value without an underlying calculation. If the value was calculated from reported data rather than copied from the study, verify the inputs and calculation and distinguish the derived result from an originally reported statistic.

Can AI report the wrong sample size from a real paper?

Yes. Papers may contain recruited, randomized, baseline, follow-up, subgroup, and analytic sample sizes. A model can select a genuine number but attach the wrong meaning to it.

Can AI calculate statistics correctly?

It can perform or assist with some calculations, especially when appropriate computational tools are available, but correctness should be verified. Check the inputs, formula or analysis procedure, assumptions, and resulting output rather than relying solely on the generated explanation.

Can AI correctly copy a statistic but still misrepresent the study?

Yes. The model may attach the statistic to the wrong outcome or interpret it incorrectly. Numerical accuracy and substantive interpretation are separate verification tasks.

Should I trust statistics when AI cites the source?

Not without checking them. A citation may be real while the number attributed to it is wrong. Conversely, the reference itself may be hallucinated. Verify both the source and the result.

Is human verification still necessary when AI extracts data for a systematic review?

Current evidence supports using LLMs as assistive rather than automatically authoritative extractors. Performance varies considerably across models, tasks, prompts, and data types, with quantitative extraction posing particular challenges in some studies. Verification procedures should therefore be built into the workflow.

09 · The Bottom Line

Every Statistical Result Needs a Verifiable Origin

The Bottom Line

Generative AI can invent statistical results, copy the wrong values, attach correct numbers to the wrong variables, or misinterpret genuine statistics.

Do not judge a number by how statistically plausible it looks. Trace consequential results to the original analysis, data, table, figure, or publication, and verify the interpretation separately. In research, a decimal point is not provenance.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes