INQUIRING LINE

People feel confident using AI, but does that confidence actually match how well the AI helps them perform?

Can self-reported confidence measures predict actual AI task performance?

This explores whether confidence, from either the people using AI or the AI models themselves, tells you how well a task will actually go, and the corpus has evidence on both sides.


This question has two halves: can people's confidence in their own AI skills predict their results, and can an AI model's confidence predict whether its answer is right? For people, the evidence says no. A pooled analysis of three studies found a correlation of only .055 between how competent people rated themselves with AI and how they actually performed, which is close enough to zero to be no signal at all Can self-ratings replace objective performance scores for AI competence?. A workplace survey shows the same gap in practice. 90% of workers felt confident using AI, but only 25% said it worked on the first try, and half spent more time with AI than they would have doing the task by hand Why do workers feel confident with AI but get poor results?. Part of the reason is that AI-assisted work makes it hard to tell whose skill produced the result. Fluent output, offloaded thinking and opaque pipelines combine so that the AI's work gets credited to the user's own competence, and each of these effects makes the others stronger How do AI tools trick users into overestimating their own skills?.

The two halves feed into each other. Users in every language studied follow how confident an AI sounds rather than whether it's accurate, so a confidently stated mistake gets followed anyway Do users worldwide trust confident AI outputs even when wrong?. When the model asserts something with confidence, the user's confidence goes up too, whether or not the answer is right. The models aren't much better judges of themselves. They can describe their own behaviors, but their self-reports are unstable and shift under conversational pressure How well do language models understand their own knowledge?. They also have no internal record of what they don't know about the user, and that gap drives both sycophancy (telling users what they want to hear) and hallucination. Simply adding a list of labeled unknowns to the prompt cut those errors by half or more Do language models know what they don't know about users?.

The surprise is that model confidence becomes useful once you stop treating it as a self-report and treat it as a measurement you check against outcomes. XConf looks up past cases where the model had a similar confidence level and reads off how often it was actually right. That matches the accuracy of sampling ten answers and comparing them, at a tenth of the cost, and the signal comes entirely from the stored track record Can past performance predict when a model will be right?. Confidence also turns out to predict something other than correctness: stability. High-confidence models give the same answer when a prompt is reworded, while low-confidence ones swing widely Does model confidence predict robustness to prompt changes?. Other work uses confidence as a steering signal instead of a verdict. It can flag when a model is overthinking or underthinking Can confidence patterns reveal overthinking versus underthinking?, and used as a training reward it can repair the calibration damage that RLHF causes (calibration meaning how well a model's confidence matches its accuracy) Can model confidence work as a reward signal for reasoning?.

The common thread is that confidence predicts performance only after it has been tied to some form of checking, and the same holds for people and machines. That links to a broader argument that AI gets good at whatever can be cheaply verified Does task verifiability determine what AI systems will learn to solve?. It also fits evaluation work where agents that gather evidence drift about 100 times less than LLMs that judge by impression Can agents evaluate AI outputs more reliably than language models?. The practical lesson is to put no weight on confidence on its own, whether it's the model's or your own, and to trust it only where something has checked it against real results.


Sources 12 notes

Can self-ratings replace objective performance scores for AI competence?

A pooled analysis of three studies found a correlation of only .055 between self-reported and objective measures of AI competence, with confidence intervals including zero. This provides no basis for substituting self-assessment for demonstrated performance.

Why do workers feel confident with AI but get poor results?

WalkMe's survey of 2,037 US workers found 90% feel confident using AI, but only 25% report it works on first try and 50% spent more time using AI than doing tasks manually. The gap widened most among younger workers, suggesting overestimation of skill.

How do AI tools trick users into overestimating their own skills?

Attribution ambiguity, fluency illusion, cognitive outsourcing, and pipeline opacity combine to systematically misattribute AI outputs as user competence. The effect is multiplicative—each mechanism amplifies the others.

Do users worldwide trust confident AI outputs even when wrong?

Cross-linguistic research shows users in every language trust confident AI outputs even when inaccurate. While confidence expression varies by language, users everywhere track confidence signals rather than accuracy, making overconfident errors systematically followed.

How well do language models understand their own knowledge?

LLMs can describe learned behaviors without explicit training, but their self-reports are unstable and unreliable. Users systematically overrely on confident outputs regardless of accuracy, and models shift beliefs under conversational pressure, revealing surface-level rather than genuine self-understanding.

Show all 12 sources
Do language models know what they don't know about users?

Research shows assistants suffer from sycophancy and hallucination because they have no representation of what remains unknown about users. Adding a schema of labeled unknowns to prompts reduced harmful advice and sycophancy by 50–75% and cut hallucination rates by roughly half.

Can past performance predict when a model will be right?

XConf matches ten-sample self-consistency at a tenth of the cost by retrieving the model's past episodes with similar confidence levels and reading their historical success rates. Ablations show the signal depends entirely on stored outcomes, not on the retrieval prompt itself.

Does model confidence predict robustness to prompt changes?

ProSA found that when models are highly confident, they resist prompt rephrasing; low confidence causes major output swings. Larger models, few-shot examples, and objective tasks all correlate with higher confidence and greater robustness.

Can confidence patterns reveal overthinking versus underthinking?

ReBalance uses confidence variance and overconfidence as diagnostic signals to apply training-free steering vectors that reduce overthinking redundancy while promoting exploration during underthinking, improving accuracy across models from 0.5B to 32B parameters.

Can model confidence work as a reward signal for reasoning?

RLSF uses answer-span confidence to rank reasoning traces, creating synthetic preferences that strengthen step-by-step reasoning while reversing RLHF's calibration degradation—without requiring human labels or external verifiers.

Does task verifiability determine what AI systems will learn to solve?

Wei argues that AI solves tasks proportional to how easily solutions can be verified, and that verifiability gaps can be narrowed by pre-investing in answer keys, test suites, or measurement infrastructure. This mechanism explains RL's effectiveness across domains from sudoku to molecular discovery.

Can agents evaluate AI outputs more reliably than language models?

Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.