Towards Accurate Differential Diagnosis with Large Language Models
An accurate differential diagnosis (DDx) is a cornerstone of medical care, often reached through an iterative process of interpretation that combines clinical history, physical examination, investigations and procedures. Interactive interfaces powered by Large Language Models (LLMs) present new opportunities to both assist and automate aspects of this process. In this study, we introduce an LLM optimized for diagnostic reasoning, and evaluate its ability to generate a DDx alone or as an aid to clinicians. 20 clinicians evaluated 302 challenging, real-world medical cases sourced from the New England Journal of Medicine (NEJM) case reports. Each case report was read by two clinicians, who were randomized to one of two assistive conditions: either assistance from search engines and standard medical resources, or LLM assistance in addition to these tools. All clinicians provided a baseline, unassisted DDx prior to using the respective assistive tools. Our LLM for DDx exhibited standalone performance that exceeded that of unassisted clinicians (top-10 accuracy 59.1% vs 33.6%, [p = 0.04]). Comparing the two assisted study arms, the DDx quality score was higher for clinicians assisted by our LLM (top-10 accuracy 51.7%) compared to clinicians without its assistance (36.1%) (McNemar’s Test: 45.7, p < 0.01) and clinicians with search (44.4%) (4.75, p = 0.03). Further, clinicians assisted by our LLM arrived at more comprehensive differential lists than those without its assistance. Our study suggests that our LLM for DDx has potential to improve clinicians’ diagnostic reasoning and accuracy in challenging cases, meriting further real-world evaluation for its ability to empower physicians and widen patients’ access to specialist-level expertise.
Introduction. An accurate diagnosis is a critical component of effective medical care. Building AI systems capable of performing or assisting clinicians in this important task has been a long-standing grand challenge [1]. While prior focus has been on evaluating a machine’s ability to accurately output a diagnosis [2–5], real-world clinical practice involves an iterative and interactive process of reasoning about a differential diagnosis (DDx), weighing multiple diagnostic possibilities in the light of increasing amounts of clinical information over time (ranging from clinical history and examination to investigations and procedures). Deep learning has been applied to promising effect for generating DDx in a number of specialties including radiology [3], ophthalmology [4] and dermatology [2], but such systems lack the interactive capabilities to fluently assist a user through communication in natural language.
The emergence of Large Language Models (LLMs) present an opportunity to design novel interactive tools and interfaces to aid in differential diagnosis. Such LLMs trained on vast corpora of text, can recognize, summarize, predict, and generate new text based on knowledge gained during the learning process and task specification via a prompt. These models have demonstrated the ability to perform complex language comprehension and reasoning tasks, generating coherent text and thereby enabling a large variety of real-world applications [6–9].
Both general-purpose LLMs (GPT-4) and medical domain-specialized LLMs (Med-PaLM 2) have demonstrated strong performance in standardized and multiple-choice medical benchmarks [10, 11]. Such evaluations represent a natural starting point for probing the medical knowledge and capabilities but fail to measure utility in real-world scenarios for care delivery, for example in challenging medical cases faced by trained physicians. It is also not obvious how these models might actively assist clinicians in the development of a DDx. Recent work has begun to assess the standalone performance of these models on challenging case reports that involve complex deduction [5, 12, 13], but has stopped short of evaluating how they can assist clinicians and augment performance and empower them to provide better care.
In this work, we introduced and investigated the ability of an LLM optimised for clinical diagnostic reasoning, to generate a DDx in challenging, real-world medical cases. Beyond measuring standalone performance like prior work [5], we integrated this model into an interactive interface to measure how well our LLM could assist clinicians in developing a DDx. Using a set of challenging real-world cases from the New England Journal of Medicine (NEJM) case reports, we compared clinicians’ ability to form a DDx with the assistance of our LLM, versus with access to traditional information retrieval tools (e.g., Internet search and books). The LLM achieved impressive performance in both generating DDx lists that contained the correct diagnosis (i.e., top-10 accuracy) and in identifying the correct final diagnosis as the most likely in the list (i.e., top-1 accuracy). Under automated model based evaluation, the quality and the accuracy of the DDx list produced by our LLM was found to be significantly better than the state-of-the-art GPT-4 model [5].
Perhaps, more importantly, the LLM also improved the diagnostic capability of clinicians as measured by the quality of their DDx lists for the evaluated cases. LLMs optimized for the safety-critical medical domain such as ours present a novel paradigm for assisting clinicians because of the potential for variation in the ways in which a given individual may converse with the system and utilise it in collaborative reasoning. We used semi-structured qualitative interviews to gather information from participating clinicians on their experiences of using the tool, their views of the potential role and risks of LLMs in medical diagnosis and in aiding the differential diagnosis process. These interviews highlighted the potential for LLMs to increase the diversity of DDx lists and speed up the process of arriving at a comprehensive DDx for challenging cases. The clinicians also highlighted that the most appropriate application at the present time would be in learning and education.
Our key contributions can be summarized as:
Method. 3 Training a Large Language Model for DDx Our study introduces an LLM for DDx, a model which uses a transformer architecture (PaLM 2 [7]), fine-tuned on medical domain data; alongside an interface for enabling its use as an interactive assistant for clinicians.
As with Med-PaLM 2 [10], our LLM builds upon PaLM 2, an iteration of Google’s LLM with substantial performance improvements on multiple LLM benchmark tasks. For the purposes of this analysis the large (L) PaLM 2 model was used.
The LLM was fine-tuned with long context length on a task mixture consisting of medical question answering (multiple-choice and long-form questions), medical dialogue generation and electronic health record (EHR) note summarization. The datasets used included the training splits of MultiMedQA (MedQA, MedMCQA, HealthSearchQA, LiveQA and MedicationQA) [10], a proprietary dataset of medical conversations, and expert handcrafted EHR note summaries from MIMIC-III [14]. The capability to process long context input enables the LLM to handle tasks that require long-range reasoning and comprehension.
Zero-Shot Prompting. We evaluated the LLM on each of the NEJM case studies with the following prompt: “You are a helpful medical assistant. You will be provided and asked about a complicated clinical case; read it carefully and then provide a diverse and thorough DDx".
4 The LLM for DDx User Interface The interface associated with our LLM, depicted in Fig. 2, enables users to interact with the underlying model via text-based chat in the context of a given case description. In our study, the interface was pre-populated with a text-only representation of the history of present illness (HPI) for a given case. Clinicians were asked to initiate the interaction by querying the LLM using a suggested prompt. Following this initial prompt and the LLM’s response, clinicians were free to query the model using any additional follow-up questions, though clinicians were cautioned to avoid asking questions about information that had not already been presented in the case. A pilot study indicated that without such a warning, clinicians may ask questions about specific lab values or imaging leading to confabulations.
In order to comparatively evaluate the LLM’s ability to generate a DDx alone and aid clinicians with their DDx generation we designed a two-stage reader study illustrated in Fig. 3). Our study was designed to evaluate the assistive effect of the LLM for generalist clinicians (not specialists) who only have access to the case presentation and not the full case information (which would include the expert commentary on the DDx). The first stage of the study had a counterbalanced design with two conditions. Clinicians generated DDx lists first without assistance and then a second time with assistance, where the type of assistance varied by condition.
Stage 1. Clinicians generate DDx with and without assistance Condition I - Search. The clinicians were first instructed to provide a list of up to ten diagnoses, with a minimum of three, based solely on review of the case presentation without using any reference materials (e.g., books) or tools (e.g., Internet search). Following this, the clinicians were instructed to use Internet Search or other resources as desired (but not given access to the LLM) and asked to re-perform their DDx.
Condition II - LLM for DDx. As with condition I, the clinicians were first instructed to provide a list of up to 10 diagnoses, with a minimum of three, based solely on review of the case presentation without using any reference materials (e.g., books) or tools (e.g., Internet search). Following this the clinicians were given access to the LLM and asked to re-perform their DDx. In addition to the LLM, clinicians could choose to use Internet search or other resources if they wished.
Stage 2. Specialists with full case information extract gold DDx and evaluate Stage 1 DDx The specialists answered the following questions to evaluate the DDx lists:
Discussion. We used a popular series of complex diagnostic challenges to evaluate an LLM optimized for clinical reasoning and diagnosis (LLM for DDx); both in a standalone capacity and under randomized comparisons as an assistive tool for physicians. In standalone performance, the LLM generated more appropriate and comprehensive DDx lists than physicians when they were unassisted, with its DDx lists more likely to include the final diagnosis than DDx lists from a board-certified internal medicine physician, no matter which position in the DDx list was considered (i.e., top-N accuracy with N ranging from 1 to 10). Clinicians using the LLM as an assistant produced a DDx with higher top-N accuracy, and DDx with greater quality, appropriateness and comprehensiveness; compared to the status quo for clinical practice (use of Internet search and other resources).
The NEJM clinicopathological conferences (CPCs) examined here are well-known for being unique and challenging clinical conundrums. Within this distinctive setting, the proposed LLM outperformed an unassisted board-certified physician in both top-1 and top-n performance. While the CPCs have long been used as benchmarks for difficult diagnosis, it is also well-known that performance in CPCs in no way reflects a broader measure of competence in a physician’s duties [16]. Furthermore, the act of DDx comprises many other steps that were not scrutinized in this study, including the goal-directed acquisition of information under uncertainty (known to be challenging for AI systems despite recent technical progress in this direction [17–19]). While Our work extends prior observations by showing not only that the LLM was more likely to arrive at a correct answer or provide the correct answer in a list, but that its DDx were determined by an independent rater to be of higher appropriateness and comprehensiveness than those produced by board certified physicians with access to references and search.
One such example is the potential for LLMs to assist clinicians in complex diagnoses. Deep learning tools have shown considerable promise in many areas of medicine, but are overwhelmingly used as assistive rather than autonomous tools [24], given the safety-critical nature of medical practice and the many issues of robustness [25] and fairness [26–28] seen in deployment. Furthermore, observations of standalone diagnostic accuracy often do not guarantee that an AI tool will improve performance in real-world settings as an assistive tool, and it remains unclear how to optimally integrate AI and human decision-making in medicine [29]. For LLMs in particular, the known incidence of hallucination/confabulation [30] might mislead clinicians into inaccurate diagnosis, replicating or even extending findings in other clinical settings that AI systems might actually degrade the performance of clinicians rather than necessarily improving outcomes.
This highlights the importance of focused study of LLMs in assistive scenarios. We explored this specifically in NEJM CPCs and found that the proposed LLM for DDx, increased the number of appropriate DDx produced by a clinician when used as an assistive tool in addition to overall top-N accuracy, suggesting that the LLM’s primary assistive potential may be due to making the scope of DDx more complete. Given the potential for misleading information to arise from AI systems, including in convincing dialogue, clinicians must appreciate the fundamental limitations of these models and not lose sight of their primacy in the provider-patient relationship and their ultimate authority and responsibility for the diagnostic and therapeutic management of their patients. Such thoughtful and effective LLM use should not be unintuitive to most clinicians.
Conclusion. Generating a DDx is a critical step in clinical case management, and the capabilities of LLMs present new opportunities for assistive tooling to help with this task. Our randomized study showed that the LLM for DDx was a helpful AI tool for DDx generation for generalist clinicians. Clinician participants indicated utility for learning and education, and additional work is needed to understand suitability for clinical settings.
Limitations. There are limitations to this evaluation. While based on real-world cases, the clinical pathology case presentation format and input into the model does differ in important ways from how a clinician would evaluate a patient and generate their differential diagnosis at the outset of a clinical encounter. The case reports are created as “puzzles” with enough clues that should enable a specialist to reason towards the final diagnosis. At the beginning of a clinician encounter, it would be challenging to create such a concise, complete and coherent case report. Case reports in the NEJM style would not be available when at intake. Similarly, these cases were selected to represent challenging cases instead of common conditions (i.e., ‘zebras’ as opposed to ‘horses’ in clinical parlance). As such, our evaluation does not directly indicate the results of or suggest that clinicians should leverage the assistive capabilities of an LLM for typical cases seen on a daily basis.
In terms of modalities, the case reports include both images and tables. The clinicians had access to these in the redacted case reports. However, the LLM only had access to the main body of the text. Though the LLM for DDx outperformed the clinicians despite this limitation, it is unknown whether and how much this gap would widen if the LLM had access to the figures and tables.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do clinicians calibrate trust in AI medical recommendations?- Can an AI system trained on text consultations handle diagnostic uncertainty in real patient encounters?
- How much does self-play training with LLM-simulated patients actually improve diagnostic accuracy?
- Does optimizing for differential diagnosis accuracy risk pushing AI systems toward premature problem-solving?
- Why do clinicians fail to act on correct AI suggestions in real care?
- How does expert annotation instability affect medical AI benchmarking?
- What evidence would prove medical AI actually works in clinics?
- Do patients actually perceive AI as worse at addressing their unique medical needs?
- How does blinded rating of diagnoses compare to real clinical outcomes?
- Does medical AI accuracy depend more on knowledge or reasoning ability?
- What prospective trials are needed to validate AI diagnostic claims?
- Does medical domain competency require knowledge injection or better prompting?
- How well do curated benchmark cases represent real clinical deployment?
- Can medical diagnosis depend less on knowledge and more on orchestration?
- How much diagnostic accuracy is gained when physicians receive expert advice?
- Does this colonoscopy finding apply to other medical specialties using AI?
- How much does prompt selection bias favor medical models over base models?
- Why did lay users with AI models fail to match unaided physicians on diagnosis?
- Do consensus criteria identify behaviors where physicians and models differ most?
- Can offline LLM evaluation predict performance in live clinical workflows?
- How much does missing images and tables limit LLM diagnostic reasoning?
- Can LLM performance on zebra cases predict results in routine clinical practice?
- Does medical fine-tuning help LLMs through knowledge or reasoning ability?
- How do LLM performances compare across different types of medical tasks?