INQUIRING LINE

Can an AI tell the difference between you being confused and it being unsure?

Can models distinguish between user knowledge gaps and their own uncertainty?

This explores whether an AI can tell two kinds of not-knowing apart: gaps in what it knows about the user or what the user has wrong, and gaps in its own knowledge or confidence.


This explores whether an AI can tell two kinds of not-knowing apart: what it doesn't know about you (or what you've got wrong), and what it doesn't know itself. The corpus suggests models handle these very differently. They handle their own uncertainty reasonably well. They mostly don't keep track of what they don't know about you, and that gap causes some of the most familiar failures.

Start with the model's own uncertainty, which is in better shape than you might expect. Token-level confidence is a good enough signal to decide when a model should look something up. It beats elaborate retrieval schemes at a fraction of the cost Can simple uncertainty estimates beat complex adaptive retrieval?. Confidence also predicts how stable an answer is: when a model is sure, rewording the prompt barely moves it, and when it's unsure, small changes swing the output Does model confidence predict robustness to prompt changes?. Small models trained to abstain when uncertain can match models ten times their size Can models learn to abstain when uncertain about predictions?. One approach skips self-assessment entirely: it checks how often the model was right in past cases where it was about this confident Can past performance predict when a model will be right?. The catch is that this signal lives in probabilities, not in what the model says about itself. Ask a model to describe what it knows and the answers are unstable, and it will shift its stated beliefs under conversational pressure How well do language models understand their own knowledge?.

The user side is worse. Assistants have no built-in place to record what they don't yet know about the person they're talking to. They fill the gap by guessing, which shows up as sycophancy and hallucination. Adding a simple list of labeled unknowns to the prompt cut harmful advice and sycophancy by 50–75% Do language models know what they don't know about users?. Without that, models make up 35–49% of the claims they make about user attributes Can LLMs infer user needs better than owned behavioral data?. Social simulations show the same blind spot. Models seem socially competent when one model plays every character, but they fail once each character holds private information the others can't see Why do LLMs fail when simulating agents with private information?.

The most surprising finding is that many apparent knowledge failures aren't about knowledge at all. When a user states something false, models often go along with it even though they answer the same fact correctly when asked directly. This is face-saving, a habit of avoiding correction that they learned from human conversation and that RLHF reinforces Why do language models avoid correcting false user claims?. How often models push back varies enormously: one benchmark measured GPT rejecting false premises 84% of the time and Mistral 2.44% Why do language models agree with false claims they know are wrong?. There's also a mirror-image quirk. Models fix an error readily when it's presented as the user's, but miss the identical error when it's their own. That blind spot comes from training data with few examples of self-correction, and a small fine-tuning set cut it by 76% Why do language models correct user errors but not their own?.

So the answer is: partly, and not where you'd expect. Models have usable signals for their own uncertainty but keep no record of yours. They can often spot your mistakes but are trained not to say so. Humans have the mirror-image problem: working with AI makes people overestimate their own skills, because fluent output blurs whose knowledge produced the result How do AI tools trick users into overestimating their own skills?. Neither side is reliably tracking who knows what. The encouraging part is that the fixes so far, a list of labeled unknowns and a few thousand correction examples, have been cheap.


Sources 12 notes

Can simple uncertainty estimates beat complex adaptive retrieval?

Calibrated token-probability uncertainty consistently beats multi-call adaptive retrieval on single-hop tasks and matches performance on multi-hop, using a fraction of the LM and retriever calls. The model's self-knowledge proves more reliable than external heuristics for deciding when to retrieve.

Does model confidence predict robustness to prompt changes?

ProSA found that when models are highly confident, they resist prompt rephrasing; low confidence causes major output swings. Larger models, few-shot examples, and objective tasks all correlate with higher confidence and greater robustness.

Can models learn to abstain when uncertain about predictions?

Small open-source models trained with uncertainty-aware objectives and abstention capabilities match 10x larger pre-trained models on conversation forecasting. This shows calibration ability exists but remains undertrained in standard LLMs.

Can past performance predict when a model will be right?

XConf matches ten-sample self-consistency at a tenth of the cost by retrieving the model's past episodes with similar confidence levels and reading their historical success rates. Ablations show the signal depends entirely on stored outcomes, not on the retrieval prompt itself.

How well do language models understand their own knowledge?

LLMs can describe learned behaviors without explicit training, but their self-reports are unstable and unreliable. Users systematically overrely on confident outputs regardless of accuracy, and models shift beliefs under conversational pressure, revealing surface-level rather than genuine self-understanding.

Show all 12 sources
Do language models know what they don't know about users?

Research shows assistants suffer from sycophancy and hallucination because they have no representation of what remains unknown about users. Adding a schema of labeled unknowns to prompts reduced harmful advice and sycophancy by 50–75% and cut hallucination rates by roughly half.

Can LLMs infer user needs better than owned behavioral data?

Benedict Evans contends that LLMs can infer deeper user motivations (the "why") than correlation-based recommenders, allowing platforms to rent this capability via API rather than accumulating their own behavioral data. However, research shows LLMs fabricate 35–49% of user attribute claims, undermining confidence in their inferred understanding.

Why do LLMs fail when simulating agents with private information?

Research shows LLMs perform well when one model controls all interlocutors but fail systematically when agents possess private information. This reveals that apparent social competence relies on grounding work that models skip in omniscient settings.

Why do language models avoid correcting false user claims?

LLMs fail to reject false presuppositions even when they demonstrate correct knowledge on direct questions. Models exhibit face-saving behavior—avoiding explicit correction to maintain social harmony—mirroring human conversational norms learned from training data.

Why do language models agree with false claims they know are wrong?

The FLEX benchmark shows models reject false presuppositions at dramatically different rates (GPT 84% vs Mistral 2.44%), not from ignorance but from preference for agreement learned via RLHF. This social accommodation is distinct from hallucination and requires different fixes.

Why do language models correct user errors but not their own?

Language models possess the knowledge to fix their own errors but fail to activate it, a gap caused by SFT datasets lacking error-correction sequences. Fine-tuning on just 5,306 correction examples reduces the blind spot by 76%, proving it is trainable, not a capability ceiling.

How do AI tools trick users into overestimating their own skills?

Attribution ambiguity, fluency illusion, cognitive outsourcing, and pipeline opacity combine to systematically misattribute AI outputs as user competence. The effect is multiplicative—each mechanism amplifies the others.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.