Strong on written medical cases, but can AI handle a real patient's messy, incomplete story and know when to say 'not sure yet'?
Can an AI system trained on text consultations handle diagnostic uncertainty in real patient encounters?
This explores whether AI that looks strong on written case vignettes and text-based consultations can cope with the messy, incomplete, uncertain information of real patients, and whether it knows when it doesn't know enough.
This explores whether AI that does well on written cases can handle real patients, where the information is partial, the patient is anxious, and the right move is often to say "I'm not sure yet." The corpus has a clear split. On curated cases, the results are very strong: one large study found an LLM beat hundreds of physicians on differential diagnosis and triage Can language models reason better than physicians at diagnosis?. Wrapping a model in an orchestration layer that decides which questions to ask and which tests to order raised accuracy on NEJM cases while cutting diagnostic costs by about 70 percent Can orchestration strategies boost diagnostic AI without better models?. That second finding is the more interesting one for uncertainty. Much of the gain came from the scaffold, not the model. In other words, managing uncertainty step by step (what to ask, what to test, when to stop) was something you could build around the model rather than something it already did well.
The real-patient evidence is thinner and more mixed. A conversational AI took histories from 100 real urgent-care patients without a single safety stop, and patients came away more positive about AI. But its management plans were less practical and less cost-aware than physicians' plans Can conversational AI safely take patient histories without supervision?. That gap fits a broader pattern. Gathering information goes fine. Acting wisely on incomplete information is harder.
The weak spot is calibration. In clinical inference tasks, LLMs pair low accuracy with high confidence, and prompting tricks that help in general domains don't reduce that overconfidence Why do language models fail confidently in specialized domains?. Patients add another pressure. Training a model to sound warm and empathetic increased its errors on medical reasoning, and the effect got worse when users expressed sadness or false beliefs Does empathy training make AI systems less reliable?. Those are the conditions of a real consultation. So the vignette results may overstate real-world performance in a specific direction: models will tend to sound more certain and more agreeable than they should.
The most useful material for building better systems comes from outside medicine. One line of work shows assistants lack any representation of what they still don't know about the user. Adding a simple list of labeled unknowns cut harmful advice and sycophancy by 50–75 percent Do language models know what they don't know about users?. Spoken dialogue systems faced the same problem decades ago with 15–30 percent speech-recognition errors. They kept a probability spread over what the user might mean instead of committing to one reading Why do dialogue systems need probabilistic reasoning?, which is close to what a differential diagnosis is. Conversation analysis gives a framework for when an agent should stop and ask a clarifying question instead of pushing ahead When should AI agents ask users instead of just searching?. And small models trained to abstain when unsure matched models ten times larger Can models learn to abstain when uncertain about predictions?. That suggests the ability to admit uncertainty is undertrained rather than missing.
A caveat: the corpus has no study that directly measures how well an AI handles diagnostic uncertainty in live encounters. The pieces have to be assembled from the work above. One reassuring finding about the human side: radiologists rated advice lower when it was labeled as AI, but their accuracy depended on whether the advice was right, not on who it came from Does labeling advice as AI change how clinicians use it?. So clinicians may be less biased against AI input than their stated opinions suggest. What the corpus points to, without being able to prove it yet, is that the hard part is not diagnosis itself but representing what is still unknown, which is a design problem more than a model-size problem.
Sources 10 notes
In physician-adjudicated vignette experiments and an emergency room study, a large language model outperformed hundreds of physicians on differential diagnosis, reasoning, triage, and clinical management tasks across multiple touchpoints.
On 304 NEJM cases, MAI-DxO orchestration achieved 79.9% accuracy at $2,397 per case versus 78.6% at $7,850 for o3 alone. The gains transferred across model families, suggesting the benefit comes from the scaffold, not model weights.
A single-arm study found that AMIE, a conversational AI system, conducted real clinical histories from 100 patients without requiring a single safety intervention by human supervisors. Patient attitudes toward AI improved after the interaction, though management plans trailed physicians on practicality and cost.
LLMs trained on general text lack sufficient exposure to domain-specific examples, leading to low accuracy paired with high confidence in clinical NLI tasks. Prompting techniques that improved general performance fail to reduce overconfidence in specialized domains.
Research shows persona training for empathy increases errors in medical reasoning, truthfulness, and disinformation resistance. Standard safety benchmarks miss this vulnerability, and effects intensify when users express sadness or false beliefs.
Show all 10 sources
Research shows assistants suffer from sycophancy and hallucination because they have no representation of what remains unknown about users. Adding a schema of labeled unknowns to prompts reduced harmful advice and sycophancy by 50–75% and cut hallucination rates by roughly half.
Real-world speech recognition achieves 15-30 percent error rates in noisy environments, making deterministic flowchart dialogue systems unworkable. POMDP-based systems handle this by maintaining belief distributions over user intent rather than committing to single interpretations.
Tool-enabled LLMs drift from user intent through silent tool chaining. Conversation analysis reveals insert-expansions—clarifying intent, scoping responses, enhancing appeal—as a formal framework for proactive user consultation that prevents misunderstanding instead of recovering from it.
Small open-source models trained with uncertainty-aware objectives and abstention capabilities match 10x larger pre-trained models on conversation forecasting. This shows calibration ability exists but remains undertrained in standard LLMs.
Radiologists rated AI-labeled advice lower than identical advice labeled human-expert, yet their diagnostic accuracy depended on whether the advice was correct, not its source. This suggests labels shape what clinicians think about advice but not how they use it.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Towards Conversational Diagnostic AI
- Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
- Sequential Diagnosis with Language Models
- Towards Accurate Differential Diagnosis with Large Language Models
- A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic
- Linguistic Calibration of Long-Form Generations
- Capabilities of Gemini Models in Medicine
- A Survey of Calibration Process for Black-Box LLMs