Clinical knowledge in LLMs does not translate to human interactions
Global healthcare providers are exploring use of large language models (LLMs) to provide medical advice to the public. LLMs now achieve nearly perfect scores on medical licensing exams, but this does not necessarily translate to accurate performance in real-world settings. We tested if LLMs can assist members of the public in identifying underlying conditions and choosing a course of action (disposition) in ten medical scenarios in a controlled study with 1,298 participants. Participants were randomly assigned to receive assistance from an LLM (GPT-4o, Llama 3, Command R+) or a source of their choice (control). Tested alone, LLMs complete the scenarios accurately, correctly identifying conditions in 94.9% of cases and disposition in 56.3% on average. However, participants using the same LLMs identified relevant conditions in less than 34.5% of cases and disposition in less than 44.2%, both no better than the control group. We identify user interactions as a challenge to the deployment of LLMs for medical advice. Standard benchmarks for medical knowledge and simulated patient interactions do not predict the failures we find with human participants. Moving forward, we recommend systematic human user testing to evaluate interactive capabilities prior to public deployments in healthcare.
Introduction. Recent breakthroughs in Artificial Intelligence research have the potential to radically transform healthcare by expanding access to medical knowledge, bringing care closer to patients. The development of large language models (LLMs) such as OpenAI’s ChatGPT could enable individuals to perform preliminary health assessments, receive personalized medical guidance, and manage chronic conditions without immediate clinician intervention. Testimonies of patients having used LLMs to successfully diagnose their own conditions are now common1. Surveys indicate that a growing number of people are already turning to AI-powered chatbots for sensitive healthrelated inquiries, with one in six American adults consulting AI chatbots for health information at least once a month2,3.
While LLMs now achieve impressive performances on medical tasks, real-world attempts to integrate LLMs into clinical settings with doctors have faced difficulties.
Research has shown that the best models can now perform on par with doctors on clinical knowledge, clinical text summarization, bedside manner, and triage—achieving scores commensurate with passing the US Medical Licensing Exam4–7. Excelling at medical question-answering tasks does not, however, translate to accurate performance in clinical settings under physician guidance. For instance, one study showed that radiologists assisted by AI did not perform better at reading chest X-rays than without AI assistance, and both performed worse than AI alone8. Another study showed that physicians assisted by LLMs only marginally outperformed unassisted physicians in diagnosis problems, and both performed worse than LLMs alone9. Providing doctors with highly capable AI systems is not enough to meaningfully assist them on important tasks10. Healthcare professionals often struggle to appropriately assess and incorporate AI-generated recommendations, limiting the benefits of AI assistance11–13.
More promising and easier to deploy, LLM-powered chatbots have instead been suggested as a ‘new front door’ to healthcare for patients who lack medical expertise14,15. As a first point of contact for healthcare support, they could be used to broaden access to medical expertise and support overburdened health systems15–17.
Medical experts have had mixed opinions on the prospects of having LLMs directly advise patients, citing problems of oversight and liability18 but also the possible benefits of providing support outside of clinical settings19,20. In response to this opportunity, private companies have made considerable efforts to create language models suitable for healthcare applications21–23.
Method. To understand if LLMs can reliably support the general public and bring care closer to patients, we conducted a randomized controlled trial with 1,298 UK participants. Each participant was tasked with identifying potential health conditions and a recommended disposition (course of action) in response to one of ten different medical scenarios. The scenarios were developed by a group of three doctors who unanimously agreed on the correct dispositions for each. The scenarios were then given to a distinct group of four doctors to provide differential diagnoses (see Fig. 1 for the study design).
We then randomly assigned participants to four experimental arms, with stratification based on demographics to ensure that each group had a composition similar to the national adult population. Participants in three treatment groups were provided with an LLM (GPT-4o, Llama 3, Command R+) for assistance in identifying conditions and dispositions. Participants in the control group were instructed to instead use any methods they would typically employ at home.
Results To assess the risks of the public using LLMs for medical advice, we conducted a randomized controlled trial where we asked participants to make decisions about a medical scenario as though they had encountered it at home (see Fig. 1). We created ten scenarios where a patient must decide whether and how to access professional medical treatment. In each scenario, participants chose the best disposition on a fivepoint scale, ranging from staying home to calling an ambulance, and listed the medical conditions they had considered which led to their choice. We scored the selected disposition based on whether it matched the answer given by the three physicians involved in drafting the case. We scored the listed medical conditions based on whether they appeared in a gold-standard list of relevant conditions generated by four physicians unfamiliar with the scenarios.
We recruited 1,298 participants living in the UK and over the age of 18 (see Fig. 1b). Participants were randomly assigned to one of three treatment groups or the control to provide a maximum of two responses, with the sample population for each experimental condition stratified to reflect the demographics of the UK. Data collection continued until 600 responses were collected for each experimental condition. Participants in the treatment groups interacted with one of GPT-4o, Llama 3, or Command R+ at least once per scenario, and as many times as they desired to help them decide how to respond to the questions. We chose these models to represent widely-used LLMs as well as the approach of using internet search to augment responses. For the control group, we instructed participants to use any assistance they would typically use at home (e.g. internet search).
Scenarios We assigned each participant two scenarios to complete consecutively. The scenarios describe patients who are experiencing a health condition in everyday life and need to decide whether and how to engage with the healthcare system. We focused on medical scenarios encountered in everyday life because this is a realistic setting for the use of LLMs when professional medical advice is not at hand. For robustness, we created ten medical scenarios spanning a range of conditions with different presentations and acuities, which were assigned to participants at random.
Discussion. This work highlights the challenges of public deployments of LLMs for direct patient care. We have conducted a randomized controlled trial of the effects of using an LLM to support medical self-assessment. Despite LLMs alone having high proficiency in the task, the combination of LLMs and human users was no better than the control group in assessing clinical acuity, and worse at identifying relevant conditions. Previous work has shown that using LLMs does not improve clinical reasoning in physicians27, and we find that this extends to the general public as well. We further identify the transmission of information between the LLM and the user as a particular point of failure, both users providing LLMs with incomplete information and LLMs suggesting correct answers but not effectively conveying this information to the users. We consider two common testing approaches for medical capabilities in LLMs, and find that although they may assess the medical information stored in the LLMs, they do not reflect the challenges of user interactions in deployment.
We specifically highlight two aspects of user interaction which impact performance in our study to motivate future research. First, over the course of the interactions, the LLMs typically offer 2-3 different possible options. This allows users to have the final decision, but they perform poorly at making this choice. Since the LLMs alone perform the task better than most users, improvements are necessary in communicating information from LLMs to users, potentially including explanations, structured outputs, or clear recommendations to help users to make decisions. The performance of the LLMs alone is a minimum estimate, since techniques such as chain of thought to use the identified conditions as a basis for the recommended disposition were not applied. Further increasing the performance of the LLMs alone would only emphasize the gap when operating with real users. Similarly, fine-tuning specialized LLMs for these tasks may be able to improve performance, but by focusing on widely-used models our results reflect the actual experience of the general public. Second, as with a real doctor-patient interaction, in this study the users have all of the information and choose what to tell the LLMs. On average, this led to weaker performance than LLMs given access to the full scenarios, though for certain scenarios, models consistently failed alone and were corrected by users. Having access to complete information is not representative of clinical practice, which indicates a need for research to develop AI systems which are designed for interaction with human users by taking into account factors such as being given incomplete or incorrect information, and the wide range of strategies users take when accessing an LLM28,29. For a public-facing medical LLM to exist, we expect that it would need to be proactive in managing and requesting information rather than relying on the user to have the expertise to guide the interaction.
With millions of people consulting LLMs for medical advice regularly2,3, healthcare practitioners will also need to know what to expect from patients with LLM-based opinions about their care. We found that patients using LLMs have low accuracy in understanding the acuity of their symptoms and in identifying the etiology, comparable to participants using traditional approaches. Our scenarios focus on common
Conclusion. As a realistic assessment of using LLMs for direct patient care, we find that current models are not ready for deployment. Overall accuracy was low, with participants using LLMs scoring below 50% on both tasks, though in a real deployment the accuracy could change depending on the relative frequencies of the scenarios. Despite strong performance from the LLMs alone, both on existing benchmarks and on our scenarios, medical expertise was insufficient for effective patient care. For policymakers and regulators considering the adoption of LLM-based systems, we recommend a focus on human user testing to evaluate interactive capabilities prior to any future deployments.
Limitations. In light of the asymmetric risks of overas opposed to underestimation, model providers may also have an incentive to push users to consult doctors rather than trust LLMs. However, we found no significant evidence that participants who consulted LLMs had higher estimates of the acuity of their scenarios, with only a small difference observed.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do clinicians calibrate trust in AI medical recommendations?- How much does self-play training with LLM-simulated patients actually improve diagnostic accuracy?
- Can an AI system trained on text consultations handle diagnostic uncertainty in real patient encounters?
- Does optimizing for differential diagnosis accuracy risk pushing AI systems toward premature problem-solving?
- Why do clinicians fail to act on correct AI suggestions in real care?
- How does expert annotation instability affect medical AI benchmarking?
- What evidence would prove medical AI actually works in clinics?
- Do patients actually perceive AI as worse at addressing their unique medical needs?
- How does blinded rating of diagnoses compare to real clinical outcomes?
- Does medical AI accuracy depend more on knowledge or reasoning ability?
- What prospective trials are needed to validate AI diagnostic claims?
- Does medical domain competency require knowledge injection or better prompting?
- How well do curated benchmark cases represent real clinical deployment?
- Can medical diagnosis depend less on knowledge and more on orchestration?
- How much diagnostic accuracy is gained when physicians receive expert advice?
- Does this colonoscopy finding apply to other medical specialties using AI?
- How much does prompt selection bias favor medical models over base models?
- Why did lay users with AI models fail to match unaided physicians on diagnosis?
- Can LLM performance on zebra cases predict results in routine clinical practice?
- Does medical fine-tuning help LLMs through knowledge or reasoning ability?
- How do LLM performances compare across different types of medical tasks?
- Can offline LLM evaluation predict performance in live clinical workflows?
- How much does missing images and tables limit LLM diagnostic reasoning?