03 · What You Need to Know
AI Confidence Has More Than One Meaning
Start by Separating Correctness From Confidence
An answer can be correct with low confidence, correct with high confidence, incorrect with low confidence, or incorrect with high confidence. Those possibilities exist for humans and machines alike.
The particularly troublesome case is the last one. An incorrect answer delivered with apparent certainty is more likely to survive superficial scrutiny because the language itself gives the user little reason to investigate further.
Researchers should therefore avoid interpreting confidence as another word for accuracy. Calibration concerns how well confidence corresponds to actual correctness across many predictions. It does not mean that every high-confidence individual answer will be correct.
Internal Confidence and Verbal Confidence Are Different
Research on language models increasingly distinguishes between internal signals associated with confidence and the confidence expressed in generated language.
A 2026 Nature Machine Intelligence study found that calibrated confidence derived from model output probabilities strongly predicted error rates and influenced whether models abstained from answering. The same research also found that verbal confidence was a separate, less discriminative signal.
This means a model can contain information relevant to how likely its answer is to be correct without perfectly translating that information into phrases such as “I am certain,” “probably,” or “I am not sure.”
Model confidence
A quantitative or internal signal associated with how strongly the model favors a prediction, which may require calibration before it can be interpreted probabilistically.
Verbal confidence
The certainty communicated through generated language, such as “definitely,” “likely,” “I am confident,” or the absence of any qualification.
Neither should be confused automatically with factual verification.
Language Models Are Trained to Produce Plausible Continuations
At the foundation of many LLMs is a next-token prediction objective. The model learns statistical patterns that allow it to predict plausible continuations of text.
This training can produce remarkably rich representations of language and knowledge. It also creates a basic asymmetry: a sentence can be linguistically probable without being factually true.
If the context strongly resembles passages in which a confident factual statement normally follows, the model can generate that form even when a particular detail is wrong. Grammatical fluency and factual verification are simply not the same objective.
Pretraining Does Not Supply a True-or-False Label for Every Claim
A 2026 Nature analysis of hallucination provides a useful explanation. Language-model pretraining typically exposes the model to enormous quantities of text but does not attach an explicit truth label to every factual statement. The model must learn factual regularities indirectly from patterns in the data.
Some patterns are highly learnable. Grammar, common expressions, and frequently repeated factual associations can become robust. Arbitrary or rare facts are more difficult because their correctness cannot necessarily be inferred from general patterns.
The result is a system that can generate the linguistic form of a correct answer even when it has failed to recover the correct fact.
Guessing Can Be Rewarded
Why not simply refuse whenever confidence is low?
One reason is that conventional evaluations often reward correct answers and count both wrong answers and abstentions as failures. Under that scoring structure, guessing can improve expected performance. The 2026 Nature analysis argues that accuracy-focused evaluation can therefore incentivize models to attempt answers rather than abstain appropriately.
This does not mean every confident error is caused by benchmarking. It does show why “always answer something” can be reinforced by the way systems are evaluated.
Researchers interested in safer behavior should therefore care not only about how often a model answers correctly, but also whether it reliably recognizes when it should not answer.
Human Preference Can Favor Convincing Explanations
Language models are often post-trained using human feedback or related preference-optimization techniques. This helps produce responses that people find useful, clear, and conversational.
Human preferences, however, are not perfectly aligned with epistemic reliability. People may prefer answers that are detailed, decisive, coherent, and easy to follow.
A study in Nature Machine Intelligence found that users systematically overestimated the accuracy of LLM answers when judging them from default explanations. Longer explanations increased user confidence even though the additional length did not improve answer accuracy.
This creates a subtle problem: characteristics that make an explanation pleasant and persuasive can also make an incorrect explanation harder to distrust.
A Wrong Answer Can Have a Perfectly Coherent Explanation
Once a model has generated an incorrect premise or answer, it can often generate a coherent explanation consistent with that answer.
Logical coherence within the response is therefore not sufficient evidence that the starting information was correct. An argument can be internally tidy while resting on a false premise, fabricated citation, incorrect number, or misunderstood study.
This is particularly dangerous in research because explanations containing disciplinary terminology can look like methodological justification rather than generated rationalization.
Confidence Can Vary by Task
LLMs are not uniformly overconfident in every setting. Calibration varies across models, tasks, prompting conditions, and confidence measures.
Some experiments find meaningful positive relationships between model confidence and correctness. For example, a Nature Human Behaviour study evaluating predictions of neuroscience results found that LLM confidence correlated positively with accuracy. Other research documents both overconfidence and underconfidence under different conditions.
The appropriate conclusion is therefore not “AI confidence is meaningless.” It is that confidence needs to be evaluated and calibrated for the task rather than inferred casually from tone.
Users Are Not Particularly Good at Detecting Incorrect Answers From Style
The communication problem is not located entirely inside the model. Humans interpret the output.
In experiments examining perceptions of LLM confidence, participants were substantially worse than model-derived confidence measures at distinguishing correct from incorrect answers based on default explanations. People tended to place too much confidence in the generated responses.
This makes intuitive sense. If you already knew enough to detect every factual error in an AI response, you would need the AI considerably less often.
The greatest risk often appears precisely when the user lacks the expertise or source access necessary to challenge a persuasive error.
Technical Detail Can Increase Persuasiveness Without Increasing Accuracy
An answer containing equations, methodological terminology, citations, effect sizes, or formal argumentation may deserve closer examination, not automatic trust.
These details can be correct and genuinely useful. They can also be generated incorrectly while retaining the surface characteristics of expertise.
This distinction leads directly to whether fluent academic-sounding AI writing indicates factual accuracy. Style provides information about how the answer is written, not independent evidence that the answer is true.
Confident Error Becomes More Dangerous When It Is Difficult to Verify
If an AI incorrectly states the capital of a country, verification takes seconds. A wrong statement about an obscure statistical assumption, a newly published study, or a highly specialized methodological debate may require substantial expertise and literature access to detect.
The cost of verification therefore varies dramatically across research tasks.
Researchers should apply stronger verification when claims are consequential, unfamiliar, highly specific, difficult to check, or dependent on information the AI may not have accessed.
Watch Out
Do not interpret the absence of uncertainty language as evidence that uncertainty is absent. A declarative sentence can be generated with the same grammatical confidence whether its content is correct, uncertain, or entirely fabricated.