01 · The Question
Can AI Give You a Statistical Result That Was Never Actually Reported?
You ask a generative AI system to summarize a study's findings. It tells you that the treatment significantly improved the outcome, perhaps adding a p-value, effect size, confidence interval, or percentage difference.
The numbers look perfectly ordinary. Nothing about p =.032 or a 95% confidence interval of 1.14 to 2.67 announces that something is wrong.
But can generative AI invent such values?
Yes. It can generate numerical results that were never reported, extract the wrong number from a real source, alter a value, confuse one statistic with another, or correctly reproduce a statistic while misinterpreting what it means. For researchers, statistical plausibility is not enough. The number needs provenance.
03 · What You Need to Know
How Can AI Get Statistical Results Wrong?
A Plausible Number Is Easy to Generate
Statistical reporting follows recognizable patterns. Research articles routinely contain means and standard deviations, percentages, coefficients, test statistics, p-values, confidence intervals, effect sizes, sample sizes, and other quantitative information.
A language model can learn those patterns. It can therefore generate a statement that looks statistically conventional without independently establishing that the number came from an actual analysis.
This is the numerical version of a broader problem with AI hallucination in research. A response can have the correct form while lacking the factual basis that the form appears to imply.
Statistically plausible
The value looks reasonable for that kind of analysis or result.
Empirically verified
The value can be traced to the actual data, analysis, or reported study result.
The first does not establish the second.
AI Can Invent a Result Entirely
The clearest failure occurs when the model provides a number for which there is no underlying analysis or source.
Suppose you ask, “What was the effect size in this study?” If the paper does not report one, an appropriate response might be that no effect size was found in the supplied information, perhaps followed by an explanation of whether one could be calculated from available data.
A hallucinated response might instead provide a value such as Cohen's d = 0.62 and describe it as a moderate effect.
The interpretation of 0.62 might be mathematically conventional. The problem is more basic: where did 0.62 come from?
Unless the value was reported or transparently calculated from verified inputs, it is not a study result.
AI Can Extract the Wrong Number From the Right Source
Numerical errors do not always involve fabrication. Scientific papers contain many numbers, often densely packed across text, tables, figures, appendices, and supplementary files.
A model may locate a genuine number but assign it to the wrong variable, group, time point, or outcome. It might report the baseline sample size as the final analytic sample, confuse a standard deviation with a standard error, or take the control group's mean as the intervention group's mean.
Recent evidence on LLM-assisted data extraction illustrates why quantitative information deserves special scrutiny. A 2026 systematic review of 27 studies evaluating LLMs for evidence-synthesis data extraction reported substantial heterogeneity in performance, with numerical data generally extracted less reliably than categorical and string variables. The review concluded that current evidence supports assistive workflows with human verification rather than autonomous extraction.
Likewise, Gougherty and Clipp found very high accuracy for several discrete and categorical ecological data-extraction tasks but substantially poorer performance for certain quantitative data, leading them to emphasize additional quality assurance.
The practical lesson is not that every AI-extracted number is suspect. It is that source access does not make numerical extraction infallible.
AI Can Change a Number Slightly and Still Change the Conclusion
Some numerical errors are visually tiny but inferentially consequential.
Consider a reported p-value of.053. If an AI summary changes it to.035, only two digits have moved. Under a conventional.05 significance threshold, however, the generated account may now describe a result as statistically significant when the reported analysis did not cross that threshold.
The same problem can occur with confidence intervals. Changing an interval from 0.94–1.21 to 1.04–1.21 may alter whether a null value is contained in the interval, depending on the effect measure.
Small transcription errors are therefore not necessarily small scientific errors.
AI Can Report a Correct Number With the Wrong Label
A numerical value is inseparable from what it represents.
A value of 0.45 could be a correlation coefficient, regression coefficient, standardized effect, probability, proportion, factor loading, or something else entirely. A number copied correctly but assigned to the wrong statistic becomes misinformation.
| What may be correct |
What AI may get wrong |
Possible consequence |
| 0.42 |
Calling a correlation an effect size from an intervention |
The study appears to demonstrate an effect it did not estimate |
| n = 312 |
Calling the recruited sample the final analytic sample |
Attrition or exclusions disappear |
| p =.041 |
Attaching it to the wrong predictor |
The wrong relationship appears statistically significant |
| 95% CI [0.70, 0.91] |
Assigning it to the wrong effect estimate |
Uncertainty is represented incorrectly |
| 68% |
Describing a subgroup percentage as the overall prevalence |
The magnitude of the finding is distorted |
AI Can Confuse Reported Statistics With Calculated Statistics
Sometimes a requested statistic is not directly reported but can legitimately be calculated from information in the paper. This creates an important distinction.
If an AI system transparently calculates a value using verified inputs and an appropriate formula, the resulting number is a derived calculation. It should be described as such.
If the system simply supplies a number and presents it as though the original authors reported it, the provenance has been misrepresented.
Reported result
A statistic explicitly presented by the original researchers.
Derived result
A statistic calculated later from reported or otherwise verified data.
Those categories should not be silently collapsed. In meta-analysis and evidence synthesis, for example, calculated effect sizes may be entirely appropriate, but the calculation method and source values need to be traceable.
Correct Statistics Can Still Receive Incorrect Interpretations
Not every statistical hallucination requires a false number. The interpretation can be wrong even when the value is copied perfectly.
An AI system might describe a statistically significant result as “important” or “large” without evidence about practical magnitude. It might treat a nonsignificant p-value as proof that there is no effect. It may describe a confidence interval as the range containing 95% of individual observations or interpret correlation as causal evidence.
Those are statistical interpretation errors rather than fabricated values, but the research consequence can be similar: the generated prose tells the reader something the analysis does not establish.
This is why AI-generated scientific claims require scrutiny even when the numbers themselves are genuine.
AI Can Generate Internally Inconsistent Statistics
Numbers in a research report are often mathematically related. A reported test statistic, p-value, confidence interval, sample size, and effect estimate may constrain one another.
A generated response can violate those relationships. It may provide a confidence interval inconsistent with the point estimate, a percentage inconsistent with the stated numerator and denominator, or a p-value that does not plausibly correspond to the reported test statistic and degrees of freedom.
Arithmetic consistency checks can catch some such errors. They cannot establish that a result is real, because a fabricated set of numbers could also be internally consistent.
Watch Out
Mathematical consistency is useful for detecting errors, but it is not provenance. A perfectly coherent set of invented statistics remains invented.
Exact Numbers Deserve More Verification, Not Less
Specificity often makes generated information feel authoritative. A statement that “the intervention improved outcomes” may invite skepticism. Add β = 0.31, p =.004, and it suddenly looks as though someone must have run an analysis.
That impression is precisely why numerical claims deserve direct verification.
The more exact the generated statistic, the easier the verification question becomes: where, exactly, does this number come from?
06 · What This Means for You
Treat Every Important AI-Generated Number as Traceable Data
Researchers do not need to manually reanalyse every statistic mentioned by AI. They do need to establish where consequential values came from.
The verification target depends on the task. If AI is summarizing a published study, compare the statistic with the relevant table, figure, results paragraph, appendix, or supplementary file. If it is analyzing your own dataset, compare the generated narrative with the actual statistical output and analysis code. If it calculated a new value, verify both the inputs and calculation.
A simple decision framework
If AI reports a statistic from a paper
Locate the exact value in the original source and confirm its variable, group, time point, and analysis.
If AI reports a statistic from your own analysis
Compare it directly with the software output, code, or verified dataset rather than the generated narrative.
If AI calculates a statistic that was not reported
Verify the input values, formula, assumptions, and calculation, and label the value as derived rather than reported.
If AI interprets a statistic
Evaluate the interpretation separately from whether the number itself was copied correctly.
If you cannot trace the number
Do not present it as an empirical result.
A useful rule is that every number important enough to appear in your research should have a route back to its origin. “The AI gave it to me” is not that route.
07 · A Quick Checklist
Before Using an AI-Generated Statistical Result
For every consequential statistic, check:
Locate the exact value in the original results, analysis output, dataset, table, figure, or supplementary material.
Confirm what statistic the number represents and which variable, group, model, or time point it belongs to.
Verify sample sizes and denominators rather than checking percentages alone.
Check signs, decimal places, confidence limits, units, and other details that can materially change interpretation.
Distinguish statistics reported by the original researchers from statistics calculated later.
Evaluate the interpretation separately from the accuracy of the numerical value.
Check related statistics for obvious internal inconsistencies, while remembering that consistency alone does not establish authenticity.
Do not publish or rely on a numerical result whose provenance you cannot establish.