INQUIRING LINE

Medical AI's right answers seem to come more from knowing the facts than reasoning well, so does math work differently?

Does medical AI accuracy depend more on knowledge or reasoning ability?

This explores whether medical AI gets answers right mainly because it knows the right facts or because it reasons well with them, and whether that balance differs from fields like math.


This explores whether medical AI's accuracy comes mainly from knowing the right medical facts or from reasoning well with them. On current evidence, knowledge matters more. When researchers scored model outputs separately for knowledge correctness and reasoning quality, medical accuracy tracked knowledge more closely, while math showed the reverse pattern Does medical AI need knowledge or reasoning more?. A practical result follows. Reasoning models distilled from DeepSeek-R1, which are strong at math, did no better than their base models on medical tasks. Better reasoning could not make up for missing or wrong medical facts, and fine-tuning did not close the gap without domain-specific data Why doesn't mathematical reasoning transfer to medicine?.

One explanation points to where these abilities sit inside the model. Knowledge retrieval seems to happen in the lower layers of the network and reasoning adjustments in the higher ones. Training that sharpens reasoning can therefore help math while leaving medical recall unchanged, or even making it worse Why does reasoning training help math but hurt medical tasks?. This fits a wider finding about how models learn. Reasoning draws on broad, general how-to knowledge that transfers across subjects, while recalling a specific fact depends on having memorized that fact from particular documents Does procedural knowledge drive reasoning more than factual retrieval?. Medicine needs many of these specific facts, and no amount of general reasoning skill supplies them.

One caution: a correct answer does not prove the reasoning behind it was good. Supervised fine-tuning can raise benchmark scores while cutting the quality of reasoning steps by 38.9 percent. The model arrives at the right answer and then writes a plausible justification after the fact Does supervised fine-tuning improve reasoning or just answers?. In medicine, a model with good recall can look like a careful clinician without reasoning like one.

The picture gets more complicated when you look at real diagnosis rather than benchmark questions. An LLM outperformed hundreds of physicians on differential diagnosis, triage and clinical management, including in an emergency room study Can language models reason better than physicians at diagnosis?. The most surprising result is about structure around the model rather than the model itself. Wrapping o3 in an orchestration framework that manages the diagnostic process step by step raised accuracy on 304 NEJM cases and cut cost per case by about 70 percent. The gains carried over to other model families, so they came from the process design, not from the model's weights Can orchestration strategies boost diagnostic AI without better models?. A related approach has the model keep an explicit list of what it doesn't yet know about a person. In a general-assistant setting, this halved hallucination and cut sycophancy and harmful advice by 50 to 75 percent Do language models know what they don't know about users?. That result comes from assistants generally, not clinical tests, but the parallel to taking a patient history is clear.

The overall picture is that inside the model, medical accuracy depends on knowledge. Clinical reasoning, though, isn't only step-by-step logic. It also means deciding what to ask next, which test is worth paying for, and what is still unknown. That kind of reasoning can be supplied by the system built around the model, which may be why scaffolding helps in medicine when reasoning training does not.


Sources 8 notes

Does medical AI need knowledge or reasoning more?

The KI/InfoGain framework reveals that medical domain accuracy correlates more strongly with knowledge correctness than reasoning quality, while mathematical domains show the inverse pattern. This distinction has direct implications for which training strategies to prioritize in each domain.

Why doesn't mathematical reasoning transfer to medicine?

R1-distilled reasoning models fail to outperform base models on medical tasks because knowledge accuracy matters more than reasoning quality in medicine—the opposite of math. Fine-tuning cannot close this gap without domain-specific training data.

Why does reasoning training help math but hurt medical tasks?

Two-phase inference model shows knowledge retrieval operates in lower network layers while reasoning adjustment happens in higher layers. This separation explains why reasoning training improves math but can degrade knowledge-intensive domains like medicine.

Does procedural knowledge drive reasoning more than factual retrieval?

Analysis of 5 million pretraining documents shows reasoning relies on broad, transferable procedural knowledge from diverse sources, unlike factual recall which depends on narrow, document-specific memorization of target facts.

Does supervised fine-tuning improve reasoning or just answers?

Supervised fine-tuning improves final-answer accuracy on benchmarks but cuts Information Gain by 38.9 percent, meaning models generate correct answers through post-hoc rationalization rather than genuine inferential steps. Standard metrics miss this degradation because they only measure final correctness.

Show all 8 sources
Can language models reason better than physicians at diagnosis?

In physician-adjudicated vignette experiments and an emergency room study, a large language model outperformed hundreds of physicians on differential diagnosis, reasoning, triage, and clinical management tasks across multiple touchpoints.

Can orchestration strategies boost diagnostic AI without better models?

On 304 NEJM cases, MAI-DxO orchestration achieved 79.9% accuracy at $2,397 per case versus 78.6% at $7,850 for o3 alone. The gains transferred across model families, suggesting the benefit comes from the scaffold, not model weights.

Do language models know what they don't know about users?

Research shows assistants suffer from sycophancy and hallucination because they have no representation of what remains unknown about users. Adding a schema of labeled unknowns to prompts reduced harmful advice and sycophancy by 50–75% and cut hallucination rates by roughly half.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.