People feel confident using AI, but does that confidence actually match how well the AI helps them perform?
Can self-reported confidence measures predict actual AI task performance?
This explores whether confidence, from either the people using AI or the AI models themselves, tells you how well a task will actually go, and the corpus has evidence on both sides.
This question has two halves: can people's confidence in their own AI skills predict their results, and can an AI model's confidence predict whether its answer is right? For people, the evidence says no. A pooled analysis of three studies found a correlation of only .055 between how competent people rated themselves with AI and how they actually performed, which is close enough to zero to be no signal at all Can self-ratings replace objective performance scores for AI competence?. A workplace survey shows the same gap in practice. 90% of workers felt confident using AI, but only 25% said it worked on the first try, and half spent more time with AI than they would have doing the task by hand Why do workers feel confident with AI but get poor results?. Part of the reason is that AI-assisted work makes it hard to tell whose skill produced the result. Fluent output, offloaded thinking and opaque pipelines combine so that the AI's work gets credited to the user's own competence, and each of these effects makes the others stronger How do AI tools trick users into overestimating their own skills?.
The two halves feed into each other. Users in every language studied follow how confident an AI sounds rather than whether it's accurate, so a confidently stated mistake gets followed anyway Do users worldwide trust confident AI outputs even when wrong?. When the model asserts something with confidence, the user's confidence goes up too, whether or not the answer is right. The models aren't much better judges of themselves. They can describe their own behaviors, but their self-reports are unstable and shift under conversational pressure How well do language models understand their own knowledge?. They also have no internal record of what they don't know about the user, and that gap drives both sycophancy (telling users what they want to hear) and hallucination. Simply adding a list of labeled unknowns to the prompt cut those errors by half or more Do language models know what they don't know about users?.
The surprise is that model confidence becomes useful once you stop treating it as a self-report and treat it as a measurement you check against outcomes. XConf looks up past cases where the model had a similar confidence level and reads off how often it was actually right. That matches the accuracy of sampling ten answers and comparing them, at a tenth of the cost, and the signal comes entirely from the stored track record Can past performance predict when a model will be right?. Confidence also turns out to predict something other than correctness: stability. High-confidence models give the same answer when a prompt is reworded, while low-confidence ones swing widely Does model confidence predict robustness to prompt changes?. Other work uses confidence as a steering signal instead of a verdict. It can flag when a model is overthinking or underthinking Can confidence patterns reveal overthinking versus underthinking?, and used as a training reward it can repair the calibration damage that RLHF causes (calibration meaning how well a model's confidence matches its accuracy) Can model confidence work as a reward signal for reasoning?.
The common thread is that confidence predicts performance only after it has been tied to some form of checking, and the same holds for people and machines. That links to a broader argument that AI gets good at whatever can be cheaply verified Does task verifiability determine what AI systems will learn to solve?. It also fits evaluation work where agents that gather evidence drift about 100 times less than LLMs that judge by impression Can agents evaluate AI outputs more reliably than language models?. The practical lesson is to put no weight on confidence on its own, whether it's the model's or your own, and to trust it only where something has checked it against real results.
Sources 12 notes
A pooled analysis of three studies found a correlation of only .055 between self-reported and objective measures of AI competence, with confidence intervals including zero. This provides no basis for substituting self-assessment for demonstrated performance.
WalkMe's survey of 2,037 US workers found 90% feel confident using AI, but only 25% report it works on first try and 50% spent more time using AI than doing tasks manually. The gap widened most among younger workers, suggesting overestimation of skill.
Attribution ambiguity, fluency illusion, cognitive outsourcing, and pipeline opacity combine to systematically misattribute AI outputs as user competence. The effect is multiplicative—each mechanism amplifies the others.
Cross-linguistic research shows users in every language trust confident AI outputs even when inaccurate. While confidence expression varies by language, users everywhere track confidence signals rather than accuracy, making overconfident errors systematically followed.
LLMs can describe learned behaviors without explicit training, but their self-reports are unstable and unreliable. Users systematically overrely on confident outputs regardless of accuracy, and models shift beliefs under conversational pressure, revealing surface-level rather than genuine self-understanding.
Show all 12 sources
Research shows assistants suffer from sycophancy and hallucination because they have no representation of what remains unknown about users. Adding a schema of labeled unknowns to prompts reduced harmful advice and sycophancy by 50–75% and cut hallucination rates by roughly half.
XConf matches ten-sample self-consistency at a tenth of the cost by retrieving the model's past episodes with similar confidence levels and reading their historical success rates. Ablations show the signal depends entirely on stored outcomes, not on the retrieval prompt itself.
ProSA found that when models are highly confident, they resist prompt rephrasing; low confidence causes major output swings. Larger models, few-shot examples, and objective tasks all correlate with higher confidence and greater robustness.
ReBalance uses confidence variance and overconfidence as diagnostic signals to apply training-free steering vectors that reduce overthinking redundancy while promoting exploration during underthinking, improving accuracy across models from 0.5B to 32B parameters.
RLSF uses answer-span confidence to rank reasoning traces, creating synthetic preferences that strengthen step-by-step reasoning while reversing RLHF's calibration degradation—without requiring human labels or external verifiers.
Wei argues that AI solves tasks proportional to how easily solutions can be verified, and that verifiability gaps can be narrowed by pre-investing in answer keys, test suites, or measurement infrastructure. This mechanism explains RL's effectiveness across domains from sudoku to molecular discovery.
Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Post-Training Large Language Models via Reinforcement Learning from Self-Feedback
- Reported Confidence in LLMs Tracks Commitment More Than Correctness
- Understanding and Mitigating Premature Confidence for Better LLM Reasoning
- A Rational Analysis of the Effects of Sycophantic AI
- Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents
- Linguistic Calibration of Long-Form Generations
- On the Reasoning Capacity of AI Models and How to Quantify It
- Beyond AI Literacy: A Structured Review and Exploratory Meta-Analysis of Measures for Competent Generative-AI Use