03 · What You Need to Know
Knowing the Answer and Knowing That You Know It Are Different Problems
What Is a Knowledge Boundary?
An LLM's knowledge boundary refers broadly to the distinction between information or tasks for which the model can produce dependable answers and those for which its knowledge or capability is insufficient.
This boundary is not necessarily simple or fixed. A model may answer a fact correctly under one phrasing and fail under another. It may know relevant background but lack one crucial detail. It may answer after receiving additional context even though the information was not available from its parameters alone.
A 2025 ACL survey of LLM knowledge boundaries emphasizes that defining and identifying these boundaries remains an active research problem rather than a solved capability.
There Are Several Reasons an AI Might Not Be Able to Answer
“I don't know” can hide several very different situations.
| Source of uncertainty |
Example |
Useful response |
| Missing knowledge |
The model lacks reliable information about a particular fact |
Abstain or retrieve an authoritative source |
| Missing context |
The answer depends on information the user has not provided |
Ask for the missing information |
| Ambiguous question |
Several interpretations of the question are plausible |
Clarify the intended meaning |
| Insufficient evidence |
The literature itself does not support a definitive conclusion |
Preserve the uncertainty |
| Capability limitation |
The model cannot reliably perform the requested reasoning or analysis |
Acknowledge the limitation or use an appropriate tool or method |
A trustworthy research assistant should ideally distinguish among these situations. Simply generating “I'm uncertain” is less useful than identifying what information or capability is missing.
Models Do Have Signals Related to Confidence
It would be too strong to claim that language models have no indication whatsoever of when an answer is uncertain. Research has identified several signals that correlate with correctness, including token probabilities, consistency across sampled responses, internal activations, and explicit confidence estimates.
A 2026 Nature Machine Intelligence study went further by investigating whether models actually use confidence to decide whether to answer or abstain. The researchers found evidence that model abstention behavior was strongly related to internal confidence and that manipulating confidence representations causally altered abstention.
This suggests that at least some contemporary models possess internal information relevant to deciding whether to withhold an answer.
It does not follow that this information is perfectly calibrated or that a conversational statement such as “I'm 95% confident” provides a literal readout of it.
Verbal Confidence Is Not the Same as Actual Reliability
When an AI says “I'm confident” or “I'm not sure,” those phrases are generated outputs. They may correlate with the model's actual likelihood of being correct, but that relationship is imperfect.
Research comparing model confidence with human perceptions has found a calibration gap: users can become more confident in AI answers than the underlying accuracy warrants. Longer explanations can further increase user confidence without necessarily increasing correctness.
That is particularly relevant in research settings because elaborate methodological reasoning can make an uncertain answer appear epistemically stronger than a short one.
Models Can Fail to Identify Why They Are Uncertain
Recognizing that something is wrong is not identical to diagnosing the problem.
ConfuseBench, introduced at ACL 2025, examined uncertainty arising from document scarcity, limited model capability, and ambiguous queries. The researchers found that current models struggled to identify the root cause of uncertainty and tended to attribute uncertainty to ambiguity while overlooking their own capability limitations.
For researchers, this matters because different uncertainties require different remedies. If the problem is missing literature, retrieve more evidence. If the question is ambiguous, clarify it. If the model cannot reliably perform the task, rephrasing the prompt indefinitely may not solve the problem.
Why Would a Model Guess Instead of Saying “I Don't Know”?
Part of the answer lies in how models are trained and evaluated.
A 2026 Nature analysis argued that conventional accuracy-based evaluation can reward guessing. If abstention is scored as simply wrong, a system trying to maximize its expected score may benefit from attempting an answer even when uncertain.
The authors contrast this with evaluations that reward appropriate abstention and penalize confident errors. This matters because model behavior reflects not only what information it contains but also what behavior its training and evaluation procedures encourage.
A model that answers every question is not necessarily more knowledgeable than one that sometimes declines. It may simply be more willing to gamble.
Abstention Can Be Improved Deliberately
Researchers are actively developing methods that encourage models to withhold answers when their knowledge is insufficient.
Approaches include confidence thresholds, calibration, prompting, fine-tuning, internal-signal methods, collaboration among models, and specialized training objectives. A 2025 ACL study on selective abstention learning, for example, trained models to reject predictions that were misaligned with the desired knowledge distribution and reported reductions in hallucination across several question-answering datasets.
Other work has taught models to express their knowledge boundaries using internal confidence signals, again improving their ability to answer known questions while declining unknown ones.
These improvements are encouraging precisely because baseline abstention is not perfectly reliable.
There Is a Trade-Off Between Answering More and Making Fewer Errors
If an AI abstains whenever it has the slightest doubt, it may avoid many errors but become frustratingly unhelpful. If it answers almost everything, coverage improves but incorrect confident answers can increase.
This is often described as a coverage-accuracy trade-off. A stricter abstention threshold means fewer questions answered but potentially higher accuracy among those that are answered.
Research systems may reasonably prefer a different threshold from casual conversational systems. When an AI is helping extract a DOI for a literature review or interpret a consequential methodological decision, saying “I cannot verify this” can be more useful than maximizing answer coverage.
An Incorrect “I Don't Know” Is Also an Error
Abstention is not automatically good. A model may decline a question it could answer correctly. Excessive abstention can make a system less useful and can obscure genuine capabilities.
The ideal system therefore does not simply say “I don't know” more often. It abstains selectively: answering when sufficiently supported and withholding or qualifying answers when the risk of error becomes unacceptable for the task.
This is why abstention research evaluates both errors and unnecessary refusals.
Retrieval Can Turn “I Don't Know” Into “Let Me Check”
For researchers, uncertainty need not always end the interaction. If a model lacks a factual answer but can search an authoritative database or retrieve the relevant paper, the appropriate behavior may be to gather more evidence rather than stop.
This distinction is especially important for research published after the model's training cutoff. The model may not possess the information parametrically but the broader AI system may be able to retrieve it.
A useful research assistant should therefore distinguish between “I don't know and cannot establish this” and “I don't currently know, but this can be verified from an appropriate source.”
The Scientific Literature Itself May Not Know
Sometimes the model's uncertainty is not the main problem. The available evidence may genuinely be insufficient or contradictory.
In such cases, forcing the AI to answer can convert scientific uncertainty into artificial certainty. The correct response may be that the scientific literature remains uncertain or divided.
This is a different form of “not knowing.” The model may accurately understand the literature and still conclude that the question does not currently have a secure scientific answer.
Self-Reported Uncertainty Should Not Replace External Verification
If the AI says “I'm not certain,” you have useful information that should prompt verification. But if it says “I'm certain,” you still need evidence when the claim matters.
The system's confidence cannot substitute for checking the source, method, calculation, or literature on which the answer depends.
This is particularly important because the model can sound confident even when it is wrong. Calibration improves the usefulness of confidence signals, but confidence and correctness remain different concepts.
Watch Out
Do not treat “I am confident” as verification or “I don't know” as proof that no answer exists. Both are model outputs. Use them as signals that can guide your next action, not as authoritative measurements of the boundary of human knowledge.
06 · What This Means for You
Give the AI Permission to Abstain, but Verify Its Boundary Yourself
Researchers can improve AI-assisted workflows by explicitly making uncertainty an acceptable output. If every prompt demands a definitive answer, the interaction itself encourages false precision.
Ask the system to identify what is known from the supplied evidence, what is inferred, what cannot be established, and what source or information would resolve the uncertainty.
A simple decision framework
If the AI says it knows the answer
Verify consequential claims from appropriate evidence rather than treating confidence as proof.
If the AI says it is uncertain
Determine whether the problem is missing evidence, ambiguous instructions, limited capability, or genuine scientific uncertainty.
If the missing information can be retrieved
Search an authoritative source or provide the relevant material instead of asking the model to guess.
If the evidence itself is insufficient
Preserve that uncertainty rather than repeatedly prompting until the model produces a definitive answer.
If the model repeatedly answers unsupported questions
Change the workflow so that claims require source grounding or verification rather than relying on self-reported confidence alone.
A particularly useful instruction for research work is to make abstention actionable: if the answer cannot be established from the available evidence, identify what is missing and where it should be verified. That turns “I don't know” from a conversational dead end into a research step.