Can orchestration strategies boost diagnostic AI without better models?
This research explores whether structuring how models collaborate—through virtual panels, cost estimation, and ensembling—can improve medical diagnosis accuracy and efficiency beyond what individual models achieve alone.
The Sequential Diagnosis Benchmark (SDBench) turns 304 New England Journal of Medicine clinicopathological conference cases, "published between 2017 and 2025," into stepwise encounters. A diagnostic agent starts from a short abstract and must request findings from a gatekeeper model, which discloses them only when asked, then commits to a diagnosis. The agent is scored on accuracy and on the estimated cost of the tests it orders. The paper's central claim is that its orchestrator, MAI-DxO, improves on the bare model along both axes. The authors report that with OpenAI's o3 it reaches 79.9% accuracy at $2,397 per case, where o3 alone reaches 78.6% at $7,850, and that its maximum-accuracy configuration reaches 85.5% at $7,184. The abstract's "four times higher" headline compares against the physician cohort, which scores 20% at $2,963 per case. Against o3 alone, the cheaper setting adds 1.3 points of accuracy (my arithmetic from the figures above); the rest of the difference is cost. These are the authors' own measurements on their own benchmark and system.
The excerpt credits the gains to "a set of physician-inspired strategies: simulating a virtual panel of physicians with distinct roles, estimating marginal costs between diagnostic rounds, and employing model ensembling methods across model responses." Its stronger claim is that these are general-purpose. MAI-DxO "boosted the accuracy of off-the-shelf models from a variety of providers by an average of 11 percentage points," and the abstract says the gains hold across OpenAI, Gemini, Claude, Grok, DeepSeek and Llama families. The mechanism sits in the scaffold, not the weights. The gatekeeper is o4-mini with the full case file and physician-written disclosure rules. The judge is o3 applying a physician-authored rubric, with four or more on a five-point scale counting as correct.
Against the nearest notes, SDBench is the clinical version of the information problem that Can models identify what information they actually need? isolates formally. QuestBench finds that models which solve a fully specified problem still fail to name the missing variable. SDBench makes information gathering the whole task. Its introduction lists the failures that static vignettes hide: "premature diagnostic closure, indiscriminate test ordering, and anchoring on early hypotheses." These are information-gathering failures of the kind QuestBench formalizes, though the paper does not test QuestBench's categories. The excerpt also qualifies Does medical AI need knowledge or reasoning more?. That note places medical competence in factual knowledge. This excerpt points to a third lever, orchestration around the model, that moves the result without a better model. Finally, the paper's limitations section concedes a version of the distortion that Do automated benchmarks hide what frontier AI systems can really do? describes. SDBench is built from "complex, pedagogically curated NEJM CPC cases," and "the case distribution does not match that of a real-world deployment scenario." It corrects the static-vignette distortion and keeps a curated-case one.
The excerpt does not establish clinical efficacy, and the authors say so: "these results do not yet establish the clinical efficacy of MAI-DxO in real-world decision support." Three further limits bear on the numbers. The cases are rare and hard, with no healthy patients, and the authors "could not measure false positive rates," so the gains on hard cases may not carry to everyday conditions. The cost figures are estimates, and the gatekeeper can "synthesize plausible results for tests not described in the original cases," so part of the cost denominator is model-generated. The physician baseline is a benchmark-condition cohort, not a record of how those physicians practice. The defensible reading is narrow. Under this benchmark's rules, an orchestrator modeled on physician reasoning gives the same model slightly higher accuracy at a fraction of the estimated cost. Whether that helps patients is a separate question the excerpt leaves open.
Inquiring lines that read this note 15
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do clinicians calibrate trust in AI medical recommendations?- Can an AI system trained on text consultations handle diagnostic uncertainty in real patient encounters?
- Does optimizing for differential diagnosis accuracy risk pushing AI systems toward premature problem-solving?
- What role does interface design play in clinician adoption of AI tools?
- How does expert annotation instability affect medical AI benchmarking?
- What evidence would prove medical AI actually works in clinics?
- Does medical AI accuracy depend more on knowledge or reasoning ability?
- What prospective trials are needed to validate AI diagnostic claims?
- What role does cost estimation play in steering diagnostic test ordering?
- Can medical diagnosis depend less on knowledge and more on orchestration?
- How much diagnostic accuracy is gained when physicians receive expert advice?
- How much do physician scores improve when assisted by the same model?
- Why did lay users with AI models fail to match unaided physicians on diagnosis?
- Does AI change clinician cognition or just increase reliance on predictions?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can models identify what information they actually need?
When a reasoning task is missing a key piece of information, can language models recognize what's absent and ask the right clarifying question? QuestBench tests this capability directly.
QuestBench isolates the missing-information failure that SDBench's gatekeeper design tests in a clinical loop
-
Do automated benchmarks hide what frontier AI systems can really do?
Benchmarks optimize for auto-gradable, short, cheap tasks. But real AI capability emerges in long-horizon, messy, open-ended work. How much capability are we missing—or wrongly inflating—by relying on benchmark scores alone?
contrast: SDBench fixes static vignettes but keeps a curated-case distortion the paper concedes
-
Does medical AI need knowledge or reasoning more?
Medical and mathematical domains may require fundamentally different AI training priorities. If medical accuracy depends primarily on factual knowledge while math depends on reasoning quality, should we build and evaluate these systems differently?
qualifies: orchestration around the model is a lever outside the knowledge-versus-reasoning split
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Sequential Diagnosis with Language Models
- The Missing Layer of AGI: From Pattern Alchemy to Coordination Physics
- Medical Reasoning in the Era of LLMs: A Systematic Review of Enhancement Techniques and Applications
- Combining Human Expertise with Artificial Intelligence: Experimental Evidence from Radiology
- Towards Accurate Differential Diagnosis with Large Language Models
- Towards Conversational Diagnostic AI
- Automation Bias in AI-Assisted Medical Decision-Making under Time Pressure in Computational Pathology
- AI-based Clinical Decision Support for Primary Care: A Real-World Study
Original note title
orchestration around o3 raised accuracy over o3 alone on NEJM cases while cutting estimated diagnostic cost 70% — and the gain transfers across model families