Does medical AI need more medical knowledge, or just better prompts, and how much of either actually helps?
Does medical domain competency require knowledge injection or better prompting?
This asks whether medical AI gets better by adding medical knowledge to the model (pretraining, fine-tuning, retrieval) or by getting more out of what it already knows through better prompts, and what the corpus says about where each approach stops working.
This asks whether medical AI gets better by adding medical knowledge to the model or by prompting it better, and where each approach runs out. The corpus answers in a way that doesn't line up with either option. Medicine depends on knowledge more than on reasoning. But the usual ways of adding knowledge help less than people expect, and the usual prompting tricks help even less. The useful question is what kind of knowledge to add and how to add it, not whether to prompt or train.
Start with what limits performance. When researchers separate getting the facts right from reasoning well, medical accuracy tracks factual correctness, while math tracks reasoning quality Does medical AI need knowledge or reasoning more?. That is why reasoning models distilled from math-heavy training (R1-style) don't beat their base models on medical tasks Why doesn't mathematical reasoning transfer to medicine?. There's a possible explanation in how the model is built: facts seem to be stored in the lower layers of the network, and reasoning adjustments happen in the higher layers. So training a model to reason more can improve math while slightly degrading recall of facts in a field like medicine Why does reasoning training help math but hurt medical tasks?. Prompting also has a hard limit here. It can only bring out knowledge the model already has, and it can't supply anything missing from training Can prompt optimization teach models knowledge they lack?. The popular "you are an expert physician" approach doesn't reliably improve factual accuracy, and giving the model a low-knowledge persona actually makes it worse Do expert personas actually improve LLM factual accuracy?. In clinical inference tasks, prompting techniques that help elsewhere don't fix a model's tendency to be confidently wrong Why do language models fail confidently in specialized domains?.
So knowledge injection should win. Here's the surprise: when seven medical LLMs were compared fairly against their general-purpose base models, with each model's prompts tuned separately and confidence intervals reported, the medical models' advantage nearly disappeared. They won on 9.4% of tasks, down from 70.5% under the usual comparison Does medical pretraining actually improve model performance?. Much of what looked like a benefit of medical pretraining came from unfair comparisons. Large general models may already contain most of the medical knowledge that bulk pretraining on medical text would add. That fits a separate result: an LLM outperformed hundreds of physicians on diagnosis, triage, and clinical management, which suggests a lot of clinical knowledge is already there Can language models reason better than physicians at diagnosis?.
The more promising direction is structure rather than volume. StructTuning organizes training text into a domain taxonomy, so the model learns where each fact fits, much as a student learns from a textbook. It reaches half the performance of full knowledge injection with 0.3% of the data Can organizing knowledge structures beat raw training data volume?. A related argument holds that systems trained only on tacit patterns, without explicit domain rules, are harder to interpret and more fragile, and that a small amount of structured knowledge goes a long way Does refusing explicit knowledge harm AI system performance?. The practical tradeoffs among retrieval (RAG), static pretraining, swappable adapters, and prompt optimization are covered in How do knowledge injection methods trade off flexibility and cost?, which finds that combining them beats any single method.
One more finding complicates the framing. In agent systems, what works as "knowledge injection" usually isn't adding facts. An analysis of 8,135 trials found that skills worked by anchoring the model to a procedure in 65.7% of cases and by supplying missing facts in only 4.5% Do skills teach procedures or inject missing facts?. For medicine, this suggests the main gap may not be missing facts. It may be retrieving the right fact reliably and applying it consistently. That is a third option between more pretraining and better prompting, and the corpus hasn't yet tested it directly on clinical tasks.
Sources 12 notes
The KI/InfoGain framework reveals that medical domain accuracy correlates more strongly with knowledge correctness than reasoning quality, while mathematical domains show the inverse pattern. This distinction has direct implications for which training strategies to prioritize in each domain.
R1-distilled reasoning models fail to outperform base models on medical tasks because knowledge accuracy matters more than reasoning quality in medicine—the opposite of math. Fine-tuning cannot close this gap without domain-specific training data.
Two-phase inference model shows knowledge retrieval operates in lower network layers while reasoning adjustment happens in higher layers. This separation explains why reasoning training improves math but can degrade knowledge-intensive domains like medicine.
Prompting works entirely within a model's pre-existing training distribution and cannot supply domain knowledge absent from training data. This creates a hard ceiling: no prompt strategy can compensate for missing foundational knowledge, only reorganize what already exists.
Testing six models on graduate-level science and engineering questions showed in-domain expert personas had no significant impact, domain-mismatched experts produced only marginal gains, and low-knowledge personas actively hurt performance. The widely-recommended role-assignment strategy lacks reliable accuracy benefit.
Show all 12 sources
LLMs trained on general text lack sufficient exposure to domain-specific examples, leading to low accuracy paired with high confidence in clinical NLI tasks. Prompting techniques that improved general performance fail to reduce overconfidence in specialized domains.
Seven medical LLMs and two medical VLMs showed significant improvements over base models in only 9.4% and 6.3% of tasks respectively when prompts were optimized per model and confidence intervals were reported, down from 70.5% and 62.5% without these controls.
In physician-adjudicated vignette experiments and an emergency room study, a large language model outperformed hundreds of physicians on differential diagnosis, reasoning, triage, and clinical management tasks across multiple touchpoints.
StructTuning achieves 50% of full-corpus performance using only 0.3% of training data by organizing chunks into auto-generated domain taxonomies. The model learns knowledge position within conceptual structures rather than raw text patterns, matching how students learn from textbooks.
AI systems that learn exclusively from data produce uninterpretable representations, inherit statistical biases uncorrected by normative rules, and fail to generalize beyond training distributions. Structured knowledge injection at minimal corpus cost substantially improves performance.
Dynamic injection (RAG) maximizes flexibility but adds latency; static embedding is fastest but costly and inflexible; modular adapters balance efficiency with swappability; prompt optimization requires no training but only activates existing knowledge. Combining all three outperforms any single approach.
Analysis of 8,135 trials shows procedural anchoring accounts for 65.7% of skill cases versus 4.5% for knowledge injection. Skills fail when retrieved incorrectly, invoked out of context, or followed too rigidly.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Capabilities of Gemini Models in Medicine
- Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains
- Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive Survey
- Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?
- Sequential Diagnosis with Language Models
- Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need
- A Survey on Knowledge Distillation of Large Language Models
- Towards Accurate Differential Diagnosis with Large Language Models