INQUIRING LINE

When an AI diagnoses a patient, does the skill come from what it knows, or from how it's organized?

Can medical diagnosis depend less on knowledge and more on orchestration?

This explores whether better AI diagnosis comes from a model knowing more medicine, or from better structure around the model: how it asks questions, orders tests, checks its own work and splits up the job.


This explores whether medical AI improves more from structure around the model (the orchestration) than from the model knowing more medicine. The corpus gives an answer most readers won't expect: it's both, because the two do different jobs. On knowledge, the evidence is strong. Medicine is a knowledge-dominant domain. Diagnostic accuracy tracks whether the model's medical facts are correct much more closely than it tracks the quality of its reasoning, and math shows the opposite pattern Does medical AI need knowledge or reasoning more?. That's why reasoning models distilled from math-heavy training don't beat their base models on medical tasks Why doesn't mathematical reasoning transfer to medicine?. There may even be a physical explanation: facts appear to be retrieved in a model's lower layers, while reasoning adjustments happen higher up. Training the reasoning layers can therefore leave the knowledge untouched, or even degrade it Why does reasoning training help math but hurt medical tasks?.

Now the surprise. Wrapping o3 in an orchestration layer called MAI-DxO, which simulates a panel of physicians who question, order tests and challenge each other, raised accuracy on 304 hard NEJM cases only slightly, from 78.6% to 79.9%. But it cut the average cost of diagnostic workup from about $7,850 to $2,397 per case. The gains also carried over to other model families Can orchestration strategies boost diagnostic AI without better models?. Look at what that pattern means. Orchestration barely changed what the model knew. It changed how the diagnosis was carried out: which tests to order, when to stop, and when to doubt a hypothesis. In real medicine that is a large part of the work, and a part that knowledge-focused benchmarks hardly measure.

A finding from outside medicine explains why these two things can be separated. When multi-step reasoning is split between a 'decomposer' that plans and a 'solver' that executes, planning skill transfers across domains but solving skill doesn't Does separating planning from execution improve reasoning accuracy?. Map that onto diagnosis: orchestration is the transferable planning layer, and medical knowledge is the domain-bound solving layer. That would explain why the MAI-DxO scaffold worked across different models while the knowledge gap still needs domain data. Related work makes the same case more generally. Putting LLM calls inside explicit algorithms, so each step sees only the context it needs, beats asking one model to do everything Can algorithms control LLM reasoning better than LLMs alone?. And tasks that need several kinds of expertise and independent checking run into organizational limits that no single agent can overcome, however capable Do single agents always hit organizational limits?.

The line between knowledge and orchestration also blurs. In one industrial case study, expert rules written into an agent's scaffolding let non-experts produce expert-rated work, with no larger model involved Can codified expertise let non-experts match specialist output?. Orchestration can carry knowledge, not just coordinate it: clinical protocols, test-ordering rules and differential checklists can live in the harness instead of the weights. The 'knowledge vs. orchestration' split also shows up in practice. Systems like Robin pair literature-reading agents with data-analysis agents in a loop with human experimenters Can multi-agent systems guide wet-lab discovery through iterative cycles?.

One caution. A single LLM has already outperformed hundreds of physicians on diagnostic vignettes and in an emergency-room study Can language models reason better than physicians at diagnosis?, so the raw knowledge may already be good enough for many cases. That makes orchestration's real contribution easy to miss. It may not raise the accuracy ceiling much, but it makes diagnosis cheaper, more disciplined and more like how medicine is actually practised. The corpus suggests the next gains will come from orchestration, as long as the underlying knowledge is sound.


Sources 10 notes

Does medical AI need knowledge or reasoning more?

The KI/InfoGain framework reveals that medical domain accuracy correlates more strongly with knowledge correctness than reasoning quality, while mathematical domains show the inverse pattern. This distinction has direct implications for which training strategies to prioritize in each domain.

Why doesn't mathematical reasoning transfer to medicine?

R1-distilled reasoning models fail to outperform base models on medical tasks because knowledge accuracy matters more than reasoning quality in medicine—the opposite of math. Fine-tuning cannot close this gap without domain-specific training data.

Why does reasoning training help math but hurt medical tasks?

Two-phase inference model shows knowledge retrieval operates in lower network layers while reasoning adjustment happens in higher layers. This separation explains why reasoning training improves math but can degrade knowledge-intensive domains like medicine.

Can orchestration strategies boost diagnostic AI without better models?

On 304 NEJM cases, MAI-DxO orchestration achieved 79.9% accuracy at $2,397 per case versus 78.6% at $7,850 for o3 alone. The gains transferred across model families, suggesting the benefit comes from the scaffold, not model weights.

Does separating planning from execution improve reasoning accuracy?

Modular architectures with separate decomposer and solver models outperform monolithic LLMs, with decomposition ability transferring across domains while solving ability does not. The separation prevents planning-execution interference and produces more generalizable skills.

Show all 10 sources
Can algorithms control LLM reasoning better than LLMs alone?

LLM Programs embed LLMs within explicit algorithms that manage control flow and state, presenting only step-specific context to each LLM call. This information hiding addresses capability and context window limits while treating complex reasoning as modular, debuggable sub-tasks.

Do single agents always hit organizational limits?

Research shows that real-world tasks requiring heterogeneous expertise, parallel execution, and independent verification exceed what any single agent loop can organize. Graph-based system abstractions are needed to distribute intelligence across specialized agents.

Can codified expertise let non-experts match specialist output?

An industrial case study embedding domain rules and design principles into an LLM agent's scaffolding achieved 206% output-quality improvement and expert-level ratings from non-experts, bypassing the need for specialist oversight. The capability gain came from externalizing tacit expertise into structured harness components, not from model scale.

Can multi-agent systems guide wet-lab discovery through iterative cycles?

Robin coordinates literature agents (Crow, Falcon) and a bioinformatic agent (Finch) in a loop where experiments inform revised hypotheses. The system proposed ripasudil for dry AMD and used consensus analysis across 10 independent trajectories, though the wet-lab validation appears only in supplementary materials.

Can language models reason better than physicians at diagnosis?

In physician-adjudicated vignette experiments and an emergency room study, a large language model outperformed hundreds of physicians on differential diagnosis, reasoning, triage, and clinical management tasks across multiple touchpoints.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.