Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Can Generative AI Reliably Tell Researchers When It Does Not Know Something?

Generative AI can express uncertainty and abstain from answering, but it does not reliably recognize every situation in which it lacks sufficient knowledge. Researchers should treat “I know” and “I don't know” as outputs to evaluate, not infallible reports of the model's internal knowledge.

29
Can AI Reliably Admit It Does Not Know? Guide 29 of 80
01 · The Question

Will AI Tell You When It Does Not Actually Know the Answer?

Researchers can tolerate an assistant that occasionally says, “I don't know.” The more dangerous assistant is one that does not know but answers anyway.

This matters because generative AI rarely looks confused in the way a human novice might. It can produce a citation, methodological explanation, historical detail, or numerical answer in polished academic prose even when the underlying information is missing or uncertain.

So when an AI lacks the answer, can it recognize that boundary and tell you?

02 · The Short Answer

AI Can Abstain, but Its Self-Knowledge Is Not Perfectly Reliable

In Brief

Generative AI can express uncertainty, decline to answer, or request more information, but researchers should not assume that it reliably recognizes every gap in its own knowledge.

Models possess signals related to confidence and can be trained or prompted to abstain more effectively, yet calibration remains imperfect. A model may answer confidently when wrong, abstain despite knowing enough to answer, or fail to identify whether uncertainty comes from missing information, an ambiguous question, or its own limited capability.

03 · What You Need to Know

Knowing the Answer and Knowing That You Know It Are Different Problems

What Is a Knowledge Boundary?

An LLM's knowledge boundary refers broadly to the distinction between information or tasks for which the model can produce dependable answers and those for which its knowledge or capability is insufficient.

This boundary is not necessarily simple or fixed. A model may answer a fact correctly under one phrasing and fail under another. It may know relevant background but lack one crucial detail. It may answer after receiving additional context even though the information was not available from its parameters alone.

A 2025 ACL survey of LLM knowledge boundaries emphasizes that defining and identifying these boundaries remains an active research problem rather than a solved capability.

There Are Several Reasons an AI Might Not Be Able to Answer

“I don't know” can hide several very different situations.

Source of uncertainty Example Useful response
Missing knowledge The model lacks reliable information about a particular fact Abstain or retrieve an authoritative source
Missing context The answer depends on information the user has not provided Ask for the missing information
Ambiguous question Several interpretations of the question are plausible Clarify the intended meaning
Insufficient evidence The literature itself does not support a definitive conclusion Preserve the uncertainty
Capability limitation The model cannot reliably perform the requested reasoning or analysis Acknowledge the limitation or use an appropriate tool or method

A trustworthy research assistant should ideally distinguish among these situations. Simply generating “I'm uncertain” is less useful than identifying what information or capability is missing.

Models Do Have Signals Related to Confidence

It would be too strong to claim that language models have no indication whatsoever of when an answer is uncertain. Research has identified several signals that correlate with correctness, including token probabilities, consistency across sampled responses, internal activations, and explicit confidence estimates.

A 2026 Nature Machine Intelligence study went further by investigating whether models actually use confidence to decide whether to answer or abstain. The researchers found evidence that model abstention behavior was strongly related to internal confidence and that manipulating confidence representations causally altered abstention.

This suggests that at least some contemporary models possess internal information relevant to deciding whether to withhold an answer.

It does not follow that this information is perfectly calibrated or that a conversational statement such as “I'm 95% confident” provides a literal readout of it.

Verbal Confidence Is Not the Same as Actual Reliability

When an AI says “I'm confident” or “I'm not sure,” those phrases are generated outputs. They may correlate with the model's actual likelihood of being correct, but that relationship is imperfect.

Research comparing model confidence with human perceptions has found a calibration gap: users can become more confident in AI answers than the underlying accuracy warrants. Longer explanations can further increase user confidence without necessarily increasing correctness.

That is particularly relevant in research settings because elaborate methodological reasoning can make an uncertain answer appear epistemically stronger than a short one.

Models Can Fail to Identify Why They Are Uncertain

Recognizing that something is wrong is not identical to diagnosing the problem.

ConfuseBench, introduced at ACL 2025, examined uncertainty arising from document scarcity, limited model capability, and ambiguous queries. The researchers found that current models struggled to identify the root cause of uncertainty and tended to attribute uncertainty to ambiguity while overlooking their own capability limitations.

For researchers, this matters because different uncertainties require different remedies. If the problem is missing literature, retrieve more evidence. If the question is ambiguous, clarify it. If the model cannot reliably perform the task, rephrasing the prompt indefinitely may not solve the problem.

Why Would a Model Guess Instead of Saying “I Don't Know”?

Part of the answer lies in how models are trained and evaluated.

A 2026 Nature analysis argued that conventional accuracy-based evaluation can reward guessing. If abstention is scored as simply wrong, a system trying to maximize its expected score may benefit from attempting an answer even when uncertain.

The authors contrast this with evaluations that reward appropriate abstention and penalize confident errors. This matters because model behavior reflects not only what information it contains but also what behavior its training and evaluation procedures encourage.

A model that answers every question is not necessarily more knowledgeable than one that sometimes declines. It may simply be more willing to gamble.

Abstention Can Be Improved Deliberately

Researchers are actively developing methods that encourage models to withhold answers when their knowledge is insufficient.

Approaches include confidence thresholds, calibration, prompting, fine-tuning, internal-signal methods, collaboration among models, and specialized training objectives. A 2025 ACL study on selective abstention learning, for example, trained models to reject predictions that were misaligned with the desired knowledge distribution and reported reductions in hallucination across several question-answering datasets.

Other work has taught models to express their knowledge boundaries using internal confidence signals, again improving their ability to answer known questions while declining unknown ones.

These improvements are encouraging precisely because baseline abstention is not perfectly reliable.

There Is a Trade-Off Between Answering More and Making Fewer Errors

If an AI abstains whenever it has the slightest doubt, it may avoid many errors but become frustratingly unhelpful. If it answers almost everything, coverage improves but incorrect confident answers can increase.

This is often described as a coverage-accuracy trade-off. A stricter abstention threshold means fewer questions answered but potentially higher accuracy among those that are answered.

Research systems may reasonably prefer a different threshold from casual conversational systems. When an AI is helping extract a DOI for a literature review or interpret a consequential methodological decision, saying “I cannot verify this” can be more useful than maximizing answer coverage.

An Incorrect “I Don't Know” Is Also an Error

Abstention is not automatically good. A model may decline a question it could answer correctly. Excessive abstention can make a system less useful and can obscure genuine capabilities.

The ideal system therefore does not simply say “I don't know” more often. It abstains selectively: answering when sufficiently supported and withholding or qualifying answers when the risk of error becomes unacceptable for the task.

This is why abstention research evaluates both errors and unnecessary refusals.

Retrieval Can Turn “I Don't Know” Into “Let Me Check”

For researchers, uncertainty need not always end the interaction. If a model lacks a factual answer but can search an authoritative database or retrieve the relevant paper, the appropriate behavior may be to gather more evidence rather than stop.

This distinction is especially important for research published after the model's training cutoff. The model may not possess the information parametrically but the broader AI system may be able to retrieve it.

A useful research assistant should therefore distinguish between “I don't know and cannot establish this” and “I don't currently know, but this can be verified from an appropriate source.”

The Scientific Literature Itself May Not Know

Sometimes the model's uncertainty is not the main problem. The available evidence may genuinely be insufficient or contradictory.

In such cases, forcing the AI to answer can convert scientific uncertainty into artificial certainty. The correct response may be that the scientific literature remains uncertain or divided.

This is a different form of “not knowing.” The model may accurately understand the literature and still conclude that the question does not currently have a secure scientific answer.

Self-Reported Uncertainty Should Not Replace External Verification

If the AI says “I'm not certain,” you have useful information that should prompt verification. But if it says “I'm certain,” you still need evidence when the claim matters.

The system's confidence cannot substitute for checking the source, method, calculation, or literature on which the answer depends.

This is particularly important because the model can sound confident even when it is wrong. Calibration improves the usefulness of confidence signals, but confidence and correctness remain different concepts.

Watch Out

Do not treat “I am confident” as verification or “I don't know” as proof that no answer exists. Both are model outputs. Use them as signals that can guide your next action, not as authoritative measurements of the boundary of human knowledge.

04 · A Practical Example

What a Useful “I Don't Know” Looks Like in Research

Hypothetical Example

An Unverifiable Paper-Specific Question

Suppose you give an AI the title of a paywalled paper that it has not accessed and ask, “What was the exact Cronbach's alpha reported for the third scale?”

Poor response The model replies, “The third scale demonstrated good reliability, with Cronbach's α =.87.”
Why this is dangerous A value around.87 may be perfectly plausible for such a scale, but the system has no accessible evidence for that exact number.
Better response The model states that it cannot verify the coefficient from the material available and identifies the full article or relevant table as necessary evidence.
Even better when tools are available The system attempts authorized retrieval, locates the article or another legitimate source containing the information, and reports the value with traceable provenance.
Researcher's role You verify the consequential numerical detail against the source before recording or citing it.

The best behavior is not merely linguistic humility. It is recognizing the evidential gap and responding appropriately to it.

05 · What Researchers Often Get Wrong

Common Misconceptions About AI Admitting Uncertainty

Misconception

If AI Does Not Know, It Will Automatically Say So

Not reliably. Models can generate plausible answers when their knowledge is insufficient, which is one of the mechanisms behind hallucination. Abstention capability is an active area of research precisely because this behavior cannot be assumed.

Misconception

If I Tell AI “Don't Hallucinate,” It Will Never Guess

Prompting can influence behavior but does not provide a guarantee. Models may still fail to recognize that a claim is unsupported or may generate an incorrect answer despite instructions to abstain when uncertain.

Misconception

An AI Confidence Percentage Is an Objective Probability of Correctness

A verbal confidence value is generated by the model and may be miscalibrated. Unless the system and confidence method have been appropriately evaluated for the task, “90% confident” should not be interpreted literally as a nine-in-ten probability of being correct.

Misconception

A Longer Explanation Means the Model Is More Certain

Explanation length and correctness are different properties. Research has found that longer explanations can increase users' confidence in AI answers even when answer accuracy does not improve.

Misconception

Saying “I Don't Know” Is Always the Safest Behavior

Unnecessary abstention can also be undesirable. A well-calibrated system should answer when adequately supported and abstain, qualify, clarify, or retrieve more information when necessary.

Misconception

If AI Admits Uncertainty, I Can Ignore the Answer

Uncertainty can tell you what to do next. You may need a source, more context, a better-defined question, an additional analysis, or qualified human expertise. A useful uncertainty signal should redirect investigation rather than merely terminate it.

06 · What This Means for You

Give the AI Permission to Abstain, but Verify Its Boundary Yourself

Researchers can improve AI-assisted workflows by explicitly making uncertainty an acceptable output. If every prompt demands a definitive answer, the interaction itself encourages false precision.

Ask the system to identify what is known from the supplied evidence, what is inferred, what cannot be established, and what source or information would resolve the uncertainty.

A simple decision framework

If the AI says it knows the answer
Verify consequential claims from appropriate evidence rather than treating confidence as proof.
If the AI says it is uncertain
Determine whether the problem is missing evidence, ambiguous instructions, limited capability, or genuine scientific uncertainty.
If the missing information can be retrieved
Search an authoritative source or provide the relevant material instead of asking the model to guess.
If the evidence itself is insufficient
Preserve that uncertainty rather than repeatedly prompting until the model produces a definitive answer.
If the model repeatedly answers unsupported questions
Change the workflow so that claims require source grounding or verification rather than relying on self-reported confidence alone.

A particularly useful instruction for research work is to make abstention actionable: if the answer cannot be established from the available evidence, identify what is missing and where it should be verified. That turns “I don't know” from a conversational dead end into a research step.

07 · A Quick Checklist

Before Trusting an AI's Claim That It Knows or Does Not Know

When AI expresses certainty or uncertainty, check:
Is the answer grounded in an identifiable source or only in model-generated knowledge?
Could the model be missing information because of its training cutoff, retrieval coverage, or lack of document access?
Is the question itself ambiguous or missing necessary context?
Could the uncertainty reflect a limitation of the model's reasoning capability rather than missing factual knowledge?
Does the scientific literature itself permit a definitive answer?
Am I treating a verbal confidence score as though it were perfectly calibrated?
Can the missing information be verified through an authoritative source or appropriate retrieval tool?
Have I allowed the model to say that the available evidence is insufficient rather than demanding a forced answer?
08 · Frequently Asked Questions

Frequently Asked Questions About AI Saying “I Don't Know”

Can an AI know when it does not know something?

Models can exhibit confidence signals related to correctness and can use those signals to support abstention behavior. Recent research provides evidence that internal confidence can influence whether some models answer or abstain. The capability remains imperfect, however, so researchers should not equate it with infallible self-knowledge.

Why does AI sometimes make something up instead of saying “I don't know”?

Generative models are trained to predict and produce responses, and conventional evaluation can reward attempts at an answer more than abstention. When information is missing, this can contribute to plausible but unsupported generation rather than explicit ignorance.

Can I prompt AI to admit uncertainty more often?

Yes. Instructions to abstain when evidence is insufficient, distinguish sourced information from inference, and identify missing information can improve behavior. Prompting does not guarantee that every knowledge gap will be detected correctly.

Should I ask AI to give a confidence percentage?

You can use confidence estimates as one signal, but do not assume a verbal percentage is statistically calibrated. The usefulness of confidence depends on the model, elicitation method, task, and validation procedure.

Is an AI that says “I don't know” frequently more trustworthy?

Not necessarily. Excessive abstention can reflect poor capability or an overly conservative threshold. Trustworthy behavior requires selective abstention: answering when adequately supported and withholding or qualifying answers when the risk of error is too high.

What should AI do instead of guessing?

Depending on the cause of uncertainty, it can abstain, qualify the answer, ask for clarification, identify missing evidence, retrieve an authoritative source, or explain that the scientific literature itself remains inconclusive.

Can AI's uncertainty replace my own verification?

No. Appropriate uncertainty can help direct verification, but important research claims still need to be checked against suitable evidence. A model's confidence is information about the model, not a substitute for evidence about the world.

09 · The Bottom Line

A Trustworthy AI Should Sometimes Refuse to Guess

The Bottom Line

Generative AI can recognize and communicate some of its uncertainty, but researchers should not rely on it to identify every knowledge gap or to abstain whenever an answer is unsupported.

Give the system permission to say “I don't know,” ask it to identify what evidence is missing, and use retrieval or verification when the uncertainty can be resolved. Most importantly, remember that confident answers still require evidence and appropriate abstention is a capability to evaluate, not a guarantee built into fluent language generation.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes