03 · What You Need to Know
Where Do AI Hallucinations Come From?
Language Models Learn to Predict Language
At the foundation of a large language model is a prediction problem. During pretraining, the model learns statistical regularities from very large collections of text by predicting tokens from their context. A token can represent a word, part of a word, punctuation, or another textual unit depending on the model.
Through this process, the model can learn an enormous amount about language and about patterns represented in its training data. It can learn that certain concepts occur together, that particular structures resemble academic references, and that certain kinds of questions tend to be followed by particular kinds of answers.
But predicting an appropriate continuation is not identical to performing an external factual verification.
Language prediction
What sequence of tokens plausibly follows from this context?
Factual verification
Is this particular claim supported by reliable evidence in the world or in an authoritative source?
The two processes can coincide. A statistically likely completion is often factually correct because the model has learned useful regularities. They can also diverge. A continuation can look exactly like the kind of answer that belongs after a question while containing a false detail.
Some Facts Are Much Harder to Learn Than Linguistic Patterns
OpenAI researchers Kalai and colleagues provide a statistical account of why some errors arise during pretraining. Their argument distinguishes learnable patterns from facts that are effectively arbitrary or difficult to infer from patterns alone.
Spelling conventions, grammatical structures, and many recurring relationships appear systematically across text. By contrast, some facts are rare, arbitrary, or poorly represented. A person's exact birthday, an obscure publication detail, or a particular identifier may not be reliably inferable merely from surrounding statistical patterns.
This matters to researchers because scholarly information contains many such particulars: author combinations, article titles, publication years, page ranges, DOI strings, sample sizes, coefficient values, instrument names, and exact quotations. A model may know the general pattern of what such information should look like without reliably possessing the specific value required in a particular case.
Training Text Does Not Come With a Universal Truth Label
Pretraining corpora are not normally databases in which every sentence has been labeled “true” or “false.” Models learn from text containing many forms of information, including disagreement, outdated claims, errors, fiction, speculation, duplicated material, and conflicting accounts.
NIST describes confabulation, another term used for hallucination, as a natural consequence of generative models approximating statistical distributions in their training data. Statistical generation can produce accurate outputs, but it can also produce statements that are factually inaccurate or internally inconsistent.
This does not mean that a model simply copies false sentences from its training data whenever it hallucinates. Hallucinations can also be newly generated combinations. The broader point is that learning linguistic distributions does not provide an automatic truth-checking mechanism for every possible proposition a model might later construct.
A Model Can Be Asked for Information It Cannot Reliably Determine
Users can ask a language model essentially unlimited questions. Some answers may be well represented in what the model has learned. Others may involve obscure facts, events outside its available information, inaccessible documents, ambiguous questions, or details that cannot be determined from the prompt.
At that point, the model faces an important behavioral choice: express uncertainty, request more information, or attempt an answer.
A system that appropriately abstains does not hallucinate merely because it lacks the answer. The problem occurs when uncertainty becomes an unsupported assertion.
Watch Out
“The model does not know” and “the model will say it does not know” are not equivalent. A system can produce a specific-looking answer even when the available information is insufficient to justify that answer.
Guessing Can Be Rewarded More Than Saying “I Don't Know”
Hallucination is not only a pretraining problem. Kalai and colleagues argue that common evaluation practices can help explain why hallucination-like guessing persists after additional training.
Consider an evaluation in which a model receives credit for a correct answer and no credit for either an incorrect answer or an abstention. When uncertain, guessing offers some possibility of being marked correct. Abstaining does not. Across many questions, a system that attempts uncertain answers may therefore achieve a higher conventional accuracy score than one that appropriately declines to answer.
This does not mean that evaluation benchmarks single-handedly cause every hallucination. Rather, it identifies an incentive problem: if systems are rewarded mainly for producing correct answers without sufficiently distinguishing confident errors from appropriate uncertainty, optimization can favor answering when abstention would be more reliable.
Question
The model receives a factual question for which it has insufficient reliable information.
Option A
It says that it does not know or cannot determine the answer.
Option B
It generates the answer that appears most plausible.
Evaluation pressure
If only exact correctness is rewarded, a lucky guess can score while an abstention cannot.
Modern AI developers use additional training, grounding, retrieval, uncertainty estimation, and other techniques intended to reduce such behavior. Hallucination rates can therefore differ substantially across models and tasks. The underlying problem, however, cannot be understood simply as a software bug that occasionally inserts random falsehoods.
Generation Itself Introduces Choices
A language model does not generally produce an entire response in one step. It generates tokens sequentially. At each stage, possible next tokens have different probabilities, and the generation procedure determines how the next token is selected.
This sequential nature matters because early choices shape what follows. Once a generated response begins down an incorrect path, subsequent tokens can form a coherent continuation of that mistaken premise. The result may become increasingly detailed even though its foundation is wrong.
Generation settings can also affect output behavior. More stochastic sampling may increase variation, for example, but hallucination cannot be reduced to a single “temperature problem.” Deterministic generation does not guarantee factuality because the most probable continuation can itself be incorrect.
Models Can Infer Patterns That Produce Plausible Fabrications
Suppose a model has encountered countless scholarly references. It has learned that references commonly contain author names, a year, an article title, a journal title, volume and issue information, page numbers, and perhaps a DOI.
That structural knowledge is useful. It also means the model can produce something that has all the characteristics of a citation without necessarily referring to a publication that exists.
The same principle can apply to scientific explanations, methodological descriptions, quotations, and numerical results. Familiarity with the form of scholarly information does not guarantee accuracy of the instance being generated.
This helps explain why AI may combine plausible authors, titles, and journals or generate other research details that fit academic conventions remarkably well.
Context and Source Grounding Can Fail Too
Not all hallucinations arise because information is absent from a model's learned parameters. A model may be given relevant information and still fail to represent it faithfully. It can overlook part of a long context, infer something that the supplied material does not establish, confuse details, or generate a conclusion beyond the evidence provided.
That distinction becomes particularly important when researchers provide the model with documents. Supplying a real paper does not automatically guarantee a faithful summary. Likewise, giving a system access to external retrieval does not mean that every sentence in the resulting answer is supported by what was retrieved.
These are reasons why hallucinations can still occur when a real research paper is supplied and why grounding technologies should be understood as mitigation mechanisms rather than automatic guarantees.