SYNTHESIS NOTE
Topics›Domain Specialization›this note

Why do LLMs fail when users interact with them?

Standard benchmarks show LLMs excel at medical diagnosis alone, yet real users get no benefit. This explores where the breakdown happens between model capability and human decision-making.

Synthesis note · 2026-10-06 · sourced from Domain Specialization

In a randomized controlled trial with 1,298 UK participants, LLMs that score well on their own did not help the public make better decisions about everyday medical scenarios. Tested alone, GPT-4o, Llama 3 and Command R+ correctly identified conditions in 94.9% of cases and dispositions in 56.3% on average. Participants using the same models identified relevant conditions in less than 34.5% of cases and dispositions in less than 44.2%, "both no better than the control group" that used whatever they would normally use at home. The paper treats this as a failure of user interaction, not of medical knowledge: standard benchmarks and simulated patient interactions "do not predict the failures we find with human participants."

The discussion locates the breakdown in transmission. The models typically offered two or three options, which "allows users to have the final decision, but they perform poorly at making this choice." Users decided what to tell the model, and the models sometimes suggested the correct answer without conveying it effectively. The combination was "no better than the control group in assessing clinical acuity, and worse at identifying relevant conditions." The authors also note that the LLM-alone figures are a minimum, because chain-of-thought reasoning over the identified conditions was not applied, and that stronger models alone "would only emphasize the gap when operating with real users."

Against the nearest notes, this trial moves a known gap from the model into the channel between model and person. Can language models truly understand therapeutic ruptures? already shows that label agreement can conceal reliance on explicit cues; here the model-alone result is strong and the failure appears only once a user is in the loop. Can language models match therapist empathy in real conversations? confines an advantage to isolated responses, and this trial finds that similar single-turn competence does not carry through a member of the public's decision. Can LLMs actually conduct Socratic questioning in therapy? places the gap in what the model can execute; this paper places part of it in the exchange itself. The patient-side note is a contrast: Why do patients distrust medical AI systems? treats adoption barriers as attitudes, while the failures here are informational and show up in task accuracy. The excerpt reports no attitudinal measures at all.

What the excerpt does not establish is how far these results travel. It covers ten everyday scenarios drafted by three doctors, one UK population stratified to national demographics, and three models used without fine-tuning or chain-of-thought prompting. The authors concede that real-deployment accuracy "could change depending on the relative frequencies of the scenarios." The account of why users fail is the authors' reading of the interactions, not a tested mechanism. The Limitations section raises an incentive question: providers may want users to see doctors rather than trust LLMs, and the authors found "no significant evidence" that LLM users rated their scenarios as more acute. The implication, at the strength one trial allows, is that benchmark scores are weak evidence about outcomes for public users, and that human user testing before public deployment is the sensible default. The trial does not show that every interactive design would fail.

Inquiring lines that read this note 5

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do clinicians calibrate trust in AI medical recommendations? What prevents LLMs from applying their reasoning knowledge to improve outputs? How do curriculum design and feedback approaches affect model learning?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
16 direct connections · 62 in 2-hop network ·sparse cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

LLMs alone correctly identified conditions in 94.9% of cases but users with the same models identified relevant conditions in less than 34.5%