03 · What You Need to Know
Why Is Completely Eliminating Hallucinations So Difficult?
“Zero Hallucinations” Depends on What Task You Mean
The question sounds binary, but the answer changes with the task.
A tightly constrained system answering questions from a small, validated database has a very different reliability problem from a general-purpose assistant asked arbitrary questions about science, history, current events, obscure publications, private documents, ambiguous situations, and information that may not exist.
Constrained task
The allowable questions, information sources, outputs, and verification rules can be narrowly controlled.
Open-ended generation
The system may receive novel, ambiguous, underspecified, obscure, or inherently unanswerable questions across many domains.
Claims about “eliminating hallucinations” should therefore specify the scope. Near-perfect performance on a bounded benchmark does not establish hallucination-free behavior across unrestricted real-world use.
Some Questions Cannot Be Reliably Answered From Available Information
One of the most important insights is that not every question has an answer available to the system.
The relevant information may not exist, may be inaccessible, may be ambiguous, may concern an unknowable future event, or may depend on facts omitted from the prompt.
OpenAI researchers Kalai and colleagues argue that real-world accuracy cannot simply reach 100% by increasing model capability because some questions are inherently unanswerable from the available information. Their analysis also distinguishes hallucination from inevitability: a model does not have to hallucinate merely because it cannot answer. It can abstain.
Question
The user asks for a fact that cannot be established from the available evidence.
Bad outcome
The system guesses and presents the guess as fact.
Better outcome
The system says that it cannot determine the answer, requests missing information, or explains the uncertainty.
The path toward lower hallucination therefore involves improving not only what models know, but also whether they recognize when they should not answer definitively.
Eliminating Errors and Eliminating Hallucinations Are Not Exactly the Same Goal
This distinction can initially seem odd.
A system may lack the information needed to answer a question correctly. If it confidently guesses, it can hallucinate. If it appropriately says “I don't know,” the question remains unanswered, but the hallucination has been avoided.
Thus, lower hallucination rates do not necessarily require perfect knowledge.
Perfect accuracy
The system correctly answers every question.
Reliable abstention
The system answers when sufficiently supported and declines or qualifies when it cannot establish the answer.
For many high-stakes applications, the second may be more realistic and more valuable than forcing an answer to every question.
Standard Accuracy Incentives Can Reward Guessing
Why do models not simply abstain whenever uncertain?
Kalai and colleagues identify one important reason in common evaluation practices. If an evaluation rewards correct answers but treats an abstention no better than a wrong answer, guessing can improve the measured accuracy score.
A model that attempts every uncertain question occasionally guesses correctly. A model that always abstains on those same questions receives no chance of earning credit.
This creates an incentive problem: systems optimized for conventional accuracy can be rewarded for behavior that increases confident errors.
Reducing hallucinations may therefore require evaluation metrics that penalize incorrect assertions more strongly and reward appropriate expressions of uncertainty.
Better Models Can Reduce Hallucinations Without Making Them Impossible
Model capability matters. Improvements in training, reasoning, calibration, post-training, tool use, and architecture can reduce error rates.
OpenAI, for example, reported in 2025 that newer models had lower hallucination rates than earlier systems while explicitly noting that hallucinations still occurred. The broader research literature similarly treats hallucination mitigation as an ongoing problem rather than a solved property of sufficiently capable models.
This distinction matters when evaluating model-level claims.
Watch Out
“Model A hallucinates less than Model B on this evaluation” does not mean “Model A will not hallucinate on my research question.” Aggregate improvement changes probability, not the evidentiary status of the particular answer in front of you.
Retrieval Can Reduce Hallucinations Without Providing a Zero-Error Guarantee
External retrieval addresses an important limitation by giving the system access to relevant evidence rather than relying solely on internal model knowledge.
Research consistently shows that retrieval-augmented generation can improve factuality. Yet contemporary RAG research continues to study retrieval errors, source reliability, context conflict, answer faithfulness, and hallucination detection precisely because retrieval does not eliminate all failure modes.
Even a system with excellent retrieval must still decide which evidence matters and generate an answer that accurately represents it.
The detailed relationship between RAG, web search, retrieval, and hallucination is therefore best understood as risk reduction rather than immunity.
Verification Can Catch Errors After Generation
Another strategy is to separate generation from checking.
A system can produce an initial answer, retrieve evidence, evaluate whether individual claims are supported, revise unsupported statements, or abstain when verification fails. Research on RAG increasingly explores explicit faithfulness assessment, uncertainty quantification, iterative fact-checking, and evidence-grounded reasoning.
This changes the architecture from:
Simple pipeline
Question → Generate answer
toward something more like:
Generate
Produce candidate claims.
Retrieve
Find relevant evidence.
Verify
Test whether the evidence supports the claims.
Revise or abstain
Correct unsupported content or decline to assert it.
Such architectures can make systems more reliable. The verification components themselves, however, must also be evaluated. A fallible model checking another fallible output does not create certainty merely by adding another stage.
Uncertainty Calibration Matters
A useful AI system should ideally distinguish situations in which its answer is strongly supported from those in which uncertainty is high.
This is known as calibration. Broadly, confidence is well calibrated when expressed or estimated confidence corresponds meaningfully to actual correctness.
Perfect calibration is difficult, particularly across changing tasks and domains. Still, better uncertainty estimation can help systems decide when to answer, retrieve additional evidence, ask for clarification, or abstain.
This is why a system saying “I cannot verify that citation” may be behaving more reliably than one that immediately supplies a complete-looking reference.
Constrained Generation Can Reduce the Space for Hallucination
Another strategy is to restrict what the model is allowed to produce.
Instead of asking for an unrestricted prose answer, a system might require outputs to be selected from a validated set, copied from verified fields, supported by retrieved passages, or returned in a schema that can be checked automatically.
This can substantially reduce certain classes of hallucination.
But constraints shift rather than magically erase the reliability problem. The system might select the wrong validated item, retrieve the wrong passage, misunderstand the user's question, or produce a formally valid but substantively incorrect output.
Hallucination Rates Depend on What Counts as a Hallucination
Another complication is definitional.
Researchers use hallucination to describe several related phenomena, including factual errors, fabricated information, contradictions, and outputs unfaithful to provided source material. NIST uses the term confabulation and includes confidently presented erroneous content as well as outputs that diverge from prompts or other inputs.
A system could therefore eliminate one narrow form of hallucination while retaining another. For example, retrieval might dramatically reduce fabricated factual answers while source-faithfulness errors remain.
Any claim that hallucinations have been “eliminated” should specify the definition, task, evaluation, and conditions under which that claim was established.
Research Reliability Should Not Depend on Waiting for Perfect AI
Researchers do not need to resolve whether hallucinations will someday disappear before deciding how to use AI now.
A more practical question is: what workflow remains reliable even when the AI occasionally gets something wrong?
That approach treats verification as part of research design rather than as an emergency response to a defective model.
For example, a workflow can require that citations be independently resolved, quotations checked against originals, statistics traced to results, and consequential factual claims linked to authoritative evidence. Better AI then makes the workflow faster and reduces the number of errors caught by those safeguards, rather than requiring the safeguards to vanish.