Sequential Diagnosis with Language Models

Paper · arXiv 2506.22405 · Published June 27, 2025
Domain Specialization in LLMs

Artificial intelligence holds great promise for expanding access to expert medical knowledge and reasoning. However, most evaluations of language models rely on static vignettes and multiple-choice questions that fail to reflect the complexity and nuance of evidence-based medicine in real-world settings. In clinical practice, physicians iteratively formulate and revise diagnostic hypotheses, adapting each subsequent question and test to what they’ve just learned, and weigh the evolving evidence before committing to a final diagnosis. To emulate this iterative diagnostic process, we introduce the Sequential Diagnosis Benchmark, which transforms 304 diagnostically challenging New England Journal of Medicine clinicopathological conference (NEJM-CPC) cases into stepwise diagnostic encounters. A physician or AI begins with a short case abstract and must iteratively request additional details from a gatekeeper model that reveals findings only when explicitly queried. Performance is assessed not just by diagnostic accuracy but also by the cost of physician visits and tests performed. To complement the benchmark, we present the MAI Diagnostic Orchestrator (MAI- DxO), a model-agnostic orchestrator that simulates a panel of physicians, proposes likely differential diagnoses and strategically selects high-value, cost-effective tests. When paired with OpenAI’s o3 model, MAI-DxO achieves 80% diagnostic accuracy—four times higher than the 20% average of generalist physicians. MAI-DxO also reduces diagnostic costs by 20% compared to physicians, and 70% compared to off-the-shelf o3. When configured for maximum accuracy, MAI-DxO achieves 85.5% accuracy. These performance gains with MAI-DxO generalize across models from the OpenAI, Gemini, Claude, Grok, DeepSeek, and Llama families. We highlight how AI systems, when guided to think iteratively and act judiciously, can advance both diagnostic precision and cost-effectiveness in clinical care.

Introduction. Sequential diagnosis is a cornerstone of clinical reasoning, wherein physicians refine their diagnostic hypotheses step-by-step through iterative questioning and testing. Figure 1 illustrates how a diagnostician might approach a case given limited initial information, posing broad then increasingly specific questions to narrow down the differential to a likely malignancy, followed by imaging, biopsy, and specialist studies to arrive at a final diagnosis. Solving such cases demands a complementary set of skills: identifying the most informative next questions or tests, balancing marginal diagnostic yield against cost and patient burden, and recognizing when the evidence is sufficient to make a confident diagnosis.

Language models (LMs) have demonstrated impressive diagnostic capabilities, with recent studies showing top-tier performance on medical licensing exams and highly structured diagnostic vignettes (Cabral et al., 2024; Goh et al., 2024; McDuff et al., 2025; Nori et al., 2023a,b, 2024). However, these evaluations occur under artificial conditions that differ markedly from real-world clinical practice. Most diagnostic assessments present models with neatly packaged vignettes that bundle the chief complaint, history of present illness, key physical exam findings, and test results, and then ask the model to select a diagnosis from a predefined answer set. By reducing the sequential diagnosis cycle to a one-turn multiple-choice quiz, static benchmarks risk overstating model competence and obscure potential weaknesses including premature diagnostic closure, indiscriminate test ordering, and anchoring on early hypotheses.

We introduce the Sequential Diagnosis Benchmark (SDBench), an interactive framework for evaluating diagnostic agents (human or AI) through realistic sequential clinical encounters. SDBench recasts 304 New England Journal of Medicine (NEJM) clinicopathological conference (CPC) cases into stepwise diagnostic encounters in which a diagnostic agent decides which questions to ask, which tests to order, and when to commit to a final diagnosis. Information is revealed by an information Gatekeeper, a language model that serves as an oracle for the patient case. The Gatekeeper discloses specific clinical findings only when explicitly queried, and can synthesize additional case-consistent information for tests not described in the original CPC narrative. Once a final diagnosis is submitted, we evaluate its correctness against the ground truth diagnosis, and compute the cumulative estimated real world cost of all requested diagnostic tests. By measuring both diagnostic accuracy and cost, SDBench aligns with the goals of the Triple Aim (Berwick et al., 2008), which seeks high quality care delivered at sustainable cost. A cohort of U.S. and U.K. physicians with a median of 12 years of experience achieved 20% accuracy at an average cost of $2,963 per case on SDBench, underscoring the inherent difficulty of the benchmark. Off-the-shelf commercial models showed varied tradeoffs: GPT-4o achieved 49.3% accuracy at a lower cost ($2,745 per case), while o3 reached 78.6% accuracy at substantially higher cost ($7,850 per case).

We further introduce MAI Diagnostic Orchestrator (MAI-DxO), an orchestrated system co-designed with physicians that consistently outperforms both human physicians and commercial language models along the cost-accuracy Pareto frontier. Compared to off-the-shelf LMs, MAI-DxO improves diagnostic accuracy while cutting estimated medical costs by more than half, demonstrating the power of careful orchestration even atop state-of-the-art models. For instance, while the off-the-shelf o3 model achieved 78.6% accuracy at a cost of $7,850, MAI-DxO achieved 79.9% at just $2,397, or 85.5% at $7,184 (Section 4). These gains stem from a set of physician-inspired strategies: simulating a virtual panel of physicians with distinct roles, estimating marginal costs between diagnostic rounds, and employing model ensembling methods across model responses. Crucially, these techniques are general-purpose: MAI-DxO boosted the accuracy of off-the-shelf models from a variety of providers by an average of 11 percentage points.

In summary, our contributions bring AI-driven diagnosis closer to clinical utility on two key fronts. First, SDBench transcends static benchmarks by aligning with the dynamic, uncertain nature of real-world

Related work. Medical problem solving has been a longstanding field of study within the medical community. In the medical AI literature, sequential diagnosis was formalized several decades ago through normative models grounded in Bayesian probability and decision theory (Horvitz et al., 1988). This framework enabled expert-level sequential diagnostic systems in domains such as nephrology (Gorry and Barnett, 1968), pathology (Heckerman et al., 1992; Horvitz et al., 1984), and trauma care (Horvitz and Seiver, More recent work has shifted toward the application of LMs to medical challenge problems, which typically include clinical reasoning as part of a broader evaluation suite (Bedi et al., 2025a; Brin et al., 2023; Chakraborty et al., 2020; Gilson et al., 2023; Gu et al., 2021; Singhal et al., 2023). While these studies demonstrated foundational performance leaps at their time of publication, existing multiplechoice benchmarks have now become saturated, highlighting the need for more complex and realistic assessments, as well as careful end-to-end agent optimization in healthcare tasks (Bedi et al., 2025b).

To this end, there have been multiple studies, notably the Articulate Medical Intelligence Explorer (AMIE) line of work, which leveraged NEJM content as source material for challenging benchmarks. For diagnostic capability assessments, AMIE also leveraged NEJM-CPC cases; however, this line of work assessed models in a fixed ”vignette” style setting in which the case information was summarized into a compact prompt and the models were asked to make a top-10 differential diagnosis (McDuff et al., 2025). In contrast, our key differentiation was to transform the static clinical case information into the real-world evidential reasoning challenge characterized by sequential diagnosis, which assesses models on their ability to iteratively ask for information, starting from minimal information, in a cost-sensitive manner and decide when a diagnosis should be made. Of note, in a parallel paper (Tu et al., 2025) AMIE was also assessed on conversational quality dimensions, such as empathy. While these represent critical dimensions of interaction with physicians and patients, we chose to frame physicians’ and agents’ interaction with SDBench as an interaction with an “oracle” about the patient, and so primarily focused on measures of cost and diagnostic accuracy. We note that (Li et al., 2024) also tests language models on information gathering capabilities; however, this work builds on much simpler, multiple-choice USMLEstyle questions (which are a few sentences long; by contrast, NEJM CPC cases are several pages long).

Method. In order to build the Sequential Diagnosis Benchmark (SDBench), we took cases from the New England Journal of Medicine’s (NEJM) Case Challenge series. The data set spans a diverse array of clinical presentations, with final diagnoses ranging from common conditions (e.g., “Covid-19 pneumonia”) to rare disorders (e.g., “Neonatal hypoglycaemia due to a biologically active teratoma”). We collected 304 consecutive cases published between 2017 and 2025, converting each into an interactive simulation of sequential diagnostic reasoning. Each encounter begins with a brief summary of the patient and their chief complaint, for example: “A 29-year-old woman was admitted to the hospital because of sore throat and peritonsillar swelling and bleeding. Symptoms did not abate with antimicrobial therapy” (Figure 1). From that starting point, a diagnostic agent (or human physician) may take one of the following actions:

  1. Ask questions: free-text questions for history or examination details (“Has she traveled recently?”). Multiple questions are allowed.

  2. Request diagnostic tests: explicit orders for labs, imaging, or procedures (“Order a CT chest with contrast”).

  3. Diagnose: a one-time commitment to a final diagnosis (“The diagnosis is histoplasmosis.”).

The Gatekeeper agent (described in detail below) interprets each request, consults the full case file, and responds in plain language, either providing the requested information or issuing a refusal if the query is too vague or non-specific. When the Diagnostic agent chooses the ‘diagnose’ action, the Judge evaluates the proposed diagnosis for correctness, and a Cost Estimator calculates the total expense of all tests ordered. The Diagnostic Agent is evaluated along two axes: diagnostic accuracy and cumulative testing cost.

Gatekeeper. We implemented the Gatekeeper using a language model (o4-mini) with access to the full NEJM CPC case file, including the final diagnosis. Guided by physician-devised rules, the Gatekeeper discloses only information that a real-world clinician could legitimately obtain from a given query or test, such as specific test results, succinct patient-history, or physical exam findings. It explicitly refuses to provide diagnostic impressions, interpret test results, or offer hints that would be unavailable in a genuine clinical encounter. Imaging is withheld until explicitly ordered; pathognomonic findings are disclosed only when the exact confirmatory test is requested; and vague or overly broad requests trigger polite refusals. Direct questions about the patient’s history or examination return responses in clinical language, closely mirroring the information extraction task faced by physicians when reviewing a medical record. Figure 1 illustrates sample requests and responses. Through this approach, the Gatekeeper removes spoilers and hindsight bias commonly embedded in educational case write-ups.

Judging diagnoses against ground truth. Two physicians may reasonably describe the same condition using different terminology, e.g. “bacterial endocarditis” versus “infective endocarditis due to Staphylococcus aureus”, yet arrive at identical treatment decisions. To account for such variability, we introduced a Judge agent to evaluate diagnoses based on clinical substance rather than surface-form descriptions. The Judge was implemented using the o3 model prompted with a detailed, physicianauthored rubric (Table 1) designed to reflect clinical consensus, similar in spirit to Arora et al. (2025). The rubric evaluates key dimensions of diagnostic quality, including the core disease entity, etiology, anatomic site, specificity, and overall completeness, with a particular emphasis on whether the candidate diagnosis would meaningfully alter clinical management. To ensure contextual understanding, the Judge had full access to each case file during adjudication. We set a cut-off of ≥4 on a five-point Likert scale to count as a “correct” diagnosis, based on the clinical rationale that clinical management would remain largely unchanged above this threshold.

Discussion. We introduce SDBench, a benchmark that transforms 304 New England Journal of Medicine CPC cases into interactive, multi-turn diagnostic challenges. Unlike static medical benchmarks that present all information upfront, SDBench more closely mirrors real-world clinical practice: diagnosticians start with minimal information and must actively decide which questions to ask, which tests to order, and when to issue a final diagnosis, with each decision incurring realistic costs. Through careful engineering, including a Gatekeeper that can synthesize plausible results for tests not described in the original cases and a clinically validated Judge to assess diagnostic accuracy, we introduce a robust evaluation environment for sequential clinical reasoning.

Within this framework, we present MAI-DxO, a system that simulates panels of different clinical personas in order to decide which questions or tests to request. MAI-DxO significantly improved diagnostic accuracy beyond strong off-the-shelf models, while simultaneously reducing cumulative test costs in SDBench, thereby establishing a new Pareto frontier between accuracy and medical cost.

When doctors begin their careers, they face a key decision: should they become generalists, with broad knowledge across many medical areas, or specialists, with deep expertise in a narrow field? This division is necessary because medicine is too vast for any one person to master in full. To manage this complexity, healthcare systems rely on collaboration: generalists and specialists work together in clinics and hospitals, combining their diverse and complementary knowledge and decision-making skills to provide patients with the comprehensive and effective care that they need.

Today, frontier AI language models are challenging this traditional structure. These advanced systems show remarkable versatility, demonstrating both broad and deep medical understanding, and the polymathic ability to reason across specialties. In effect, they combine the generalist’s range with specialists’ depth. As a result, they significantly outperform individual physicians on complex diagnostic problems, such as those featured in the NEJM CPC cases. Our findings highlight this impressive capability. Expecting any single doctor to master the full range of such cases is unrealistic.

This raises an intriguing question: When evaluating frontier AI systems, should we evaluate frontier AI systems by comparing them to individual physicians, or to entire hospital-like teams of generalists and specialists? The answer to this question will help both define and shape the future role of AI in healthcare.

Conclusion. Our findings demonstrate the promise of AI methods for sequential diagnosis, including the ability to explicitly model working differential diagnoses and reason about informational value and cost of diagnostic tests. While these results do not yet establish the clinical efficacy of MAI-DxO in real-world decision support, they underscore AI’s increasing potential to address urgent challenges in healthcare delivery. Our model-agnostic system design may alleviate risks and implementation challenges for health systems aiming to adopt best-in-class language model–based diagnostic support in a rapidly evolving field. By reducing reliance on any single model, it avoids the need to “version chase” each new model release. In terms of practical application, future work should validate MAI-DxO in everyday clinical environments, where disease prevalence and presentations reflect routine practice rather than the rare, complex cases featured in the NEJM CPC corpus. An immediate goal is to identify the settings in which MAI-DxO could address unmet needs and deliver the greatest value to health outcomes and societal benefit.

We hypothesize that access to superhuman diagnostic capabilities requiring minimal health IT infrastructure could improve quality of care globally, helping to mitigate the costly impact of clinical workforce shortages and variability in care delivery Mandl (2025); Wennberg et al.. In resource-limited settings especially, cost-effective strategies may enable health systems to impact more lives per dollar spent, allowing scarce medical resources to be reserved for those with the most urgent clinical needs. More broadly, such systems might even make direct-to-consumer tools possible, such as smartphone-based triage, provided that safety, regulatory clearance, and data-privacy safeguards are demonstrably in place.

Progress toward effective clinical decision support will require the development of diagnostic corpora that mirror real-world prevalence patterns. Such benchmarks will help to surface limitations and opportunities for refinement that may be obscured by our current emphasis on especially difficult diagnostic scenarios.

Limitations. Since SDBench is built from complex, pedagogically curated NEJM CPC cases, the case distribution does not match that of a real-world deployment scenario, and indeed there are no cases where the patients are in fact healthy or have benign syndromes. Thus, we do not know whether MAI-DxO’s performance gains on hard cases generalize to common, everyday clinical conditions, and could not measure false positive rates. Additionally, a practical diagnostic agent must incorporate patient-specific risk factors, and consider additional factors beyond cost, e.g.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do clinicians calibrate trust in AI medical recommendations? How do curriculum design and feedback approaches affect model learning? Why do standard evaluation practices obscure safety-critical AI failures? What prevents LLMs from applying their reasoning knowledge to improve outputs? What limits language model accuracy in evaluating ideas?