INQUIRING LINE

When an AI sounds equally sure about a diagnosis it knows and one it doesn't, can you trust its confidence?

How well-calibrated are language models when making clinical predictions?

This explores whether a language model's confidence in a medical judgment matches how often it is actually right, and what the corpus says about why that match breaks down and how it might be repaired.


This explores whether a model's confidence in a clinical judgment tracks how often it's actually right. Being accurate and being calibrated are separate questions. The collection's most direct evidence is not encouraging: on clinical inference tasks, models pair low accuracy with high confidence, and prompting techniques that improve general performance do nothing to reduce that overconfidence Why do language models fail confidently in specialized domains?. The explanation offered is exposure. General web text contains too few worked clinical examples for the model to learn where its knowledge runs out, so it sounds just as sure in unfamiliar territory as in familiar territory.

The complication is that accuracy alone can look very good. In vignette experiments and an emergency-room study, one model beat hundreds of physicians on differential diagnosis, triage, and management Can language models reason better than physicians at diagnosis?. Taken together, the two notes suggest something worth knowing: a model can outperform doctors on average and still not know which of its answers to trust. In a clinic, that second skill decides when a human should step in. Strong benchmark accuracy doesn't show it.

Looking sideways at how calibration gets lost and recovered helps. One line of work argues that RLHF, the training that makes models pleasant assistants, also degrades calibration. Using the model's own answer confidence as a training reward appears to reverse that damage while improving reasoning Can model confidence work as a reward signal for reasoning?. Another approach stops trusting the model's in-the-moment sense of certainty. Instead, it retrieves the model's past cases at a similar confidence level and checks how often those turned out right Can past performance predict when a model will be right?. That idea maps neatly onto medicine, where a model could be calibrated against its own logged outcomes on real patients rather than its self-report.

There is also a quieter clinical risk. Models sometimes ignore information in front of them when it conflicts with strong associations learned in training Why do language models ignore information in their context?, and larger, instruction-tuned models do this more often Do larger models follow stated beliefs less often?. For an atypical patient whose details cut against the textbook picture, the model may answer the textbook case with textbook confidence. The 'embers of autoregression' framing adds to this: tasks whose correct answer is statistically unlikely are systematically harder, even when they're logically simple Can we predict where language models will fail?. Applying this to rare presentations is an inference, not a tested clinical result.

An honest limit: the corpus has only one note that directly measures clinical calibration. Everything else is adjacent evidence about why calibration fails and how it might be fixed. The open question these notes point toward is whether outcome-grounded confidence methods have been tried in real clinical deployments. The collection doesn't yet answer that.


Sources 7 notes

Why do language models fail confidently in specialized domains?

LLMs trained on general text lack sufficient exposure to domain-specific examples, leading to low accuracy paired with high confidence in clinical NLI tasks. Prompting techniques that improved general performance fail to reduce overconfidence in specialized domains.

Can language models reason better than physicians at diagnosis?

In physician-adjudicated vignette experiments and an emergency room study, a large language model outperformed hundreds of physicians on differential diagnosis, reasoning, triage, and clinical management tasks across multiple touchpoints.

Can model confidence work as a reward signal for reasoning?

RLSF uses answer-span confidence to rank reasoning traces, creating synthetic preferences that strengthen step-by-step reasoning while reversing RLHF's calibration degradation—without requiring human labels or external verifiers.

Can past performance predict when a model will be right?

XConf matches ten-sample self-consistency at a tenth of the cost by retrieving the model's past episodes with similar confidence levels and reading their historical success rates. Ablations show the signal depends entirely on stored outcomes, not on the retrieval prompt itself.

Why do language models ignore information in their context?

Research demonstrates that LMs generate outputs inconsistent with their context because parametric knowledge from training dominates over in-context information. Textual prompting alone cannot override strong priors; causal intervention in representations is required.

Show all 7 sources
Do larger models follow stated beliefs less often?

Across 18 LLMs tested with EoBench, bigger models and instruction-tuned variants showed lower rates of context-following when users expressed beliefs that contradicted world knowledge. The effect suggests instruction-tuning strengthens reliance on parametric knowledge.

Can we predict where language models will fail?

By framing LLMs as autoregressive probability machines, researchers predicted tasks with low-probability target responses would be systematically harder, even when logically simple. Experiments confirmed predictions like backwards alphabet and letter counting.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.