Medical AI models look better than general ones, but under fair testing they significantly beat them on under 10% of tasks.
Which medical tasks benefit most from domain-adaptive pretraining?
This explores which kinds of medical work, such as answering exam questions, clinical reasoning, or reading clinical notes, actually get better when a model is given extra training on medical text before anything else. The corpus's main answer is that fewer tasks benefit than the headline numbers suggest.
This explores which medical tasks really improve when a general model gets extra pretraining on medical text. The answer starts by questioning the premise. When seven medical LLMs were compared fairly with their own base models, each model got its own tuned prompts and the results were checked for statistical noise. Under those conditions the medical versions were significantly better on only 9.4% of tasks. Without those controls they looked better on 70.5% (Does medical pretraining actually improve model performance?). Medical vision-language models fell from 62.5% to 6.3%. The corpus doesn't name a group of medical tasks that clearly wins. Most of the apparent benefit came from comparing a well-prompted specialist against a poorly prompted generalist.
That result fits another finding. Prompting can't give a model knowledge it never learned. It can only bring out knowledge already stored somewhere in the model (Can prompt optimization teach models knowledge they lack?). So if good prompting closes most of the gap, the general models already knew most of the medicine. Medical pretraining mostly made that knowledge easier to reach. It didn't add much that was new. The places to look for real gains are therefore tasks where general training data is thin. One example is clinical inference: deciding whether one clinical statement follows from another. There, general models are often wrong and highly confident at the same time, and clever prompting doesn't reduce that overconfidence (Why do language models fail confidently in specialized domains?).
A more useful way to frame the question is to ask what kind of change a medical task needs. One line of work suggests that factual knowledge lives mostly in a model's lower layers and reasoning in its higher layers. That would explain why training a model to reason better helps with math but can hurt medicine, where the bottleneck is remembering facts correctly (Why does reasoning training help math but hurt medical tasks?). Reinforcement learning on medical reasoning seems to help mainly by getting the model to stop using wrong facts in its reasoning, not by teaching it new ones (Does RL improve domain reasoning by adding knowledge or removing it?). Other approaches reward the model for giving a sound explanation as well as the right answer, and these seem to build knowledge in more firmly than standard fine-tuning (Can reinforcement learning embed domain knowledge more effectively than supervised fine-tuning?).
There's also a hidden cost to watch for. Training directly on domain data can damage the very knowledge stored in those lower layers. Methods that adjust the model's outputs while generating text, leaving its weights untouched, keep that knowledge intact better (Can decoding-time tuning preserve knowledge better than weight fine-tuning?). Benchmark scores can also mislead. Supervised fine-tuning can raise final-answer accuracy while the reasoning behind each answer gets worse, with the model making up a justification after reaching the answer (Does supervised fine-tuning improve reasoning or just answers?). Every adaptation method has a narrow set of conditions where it works best, and its benefits often come with degradation that is hard to see (How do domain training techniques actually reshape model behavior?).
The surprising lesson is that "which medical tasks benefit?" may be less important than "was the comparison fair?" Many claims for specialist medical models fade once the general model gets the same care with prompting. Where gains do hold up, they are most likely on rare, specialized material the general model never saw. They are least likely on standard medical-exam questions that are already well covered on the open web.
Sources 9 notes
Seven medical LLMs and two medical VLMs showed significant improvements over base models in only 9.4% and 6.3% of tasks respectively when prompts were optimized per model and confidence intervals were reported, down from 70.5% and 62.5% without these controls.
Prompting works entirely within a model's pre-existing training distribution and cannot supply domain knowledge absent from training data. This creates a hard ceiling: no prompt strategy can compensate for missing foundational knowledge, only reorganize what already exists.
LLMs trained on general text lack sufficient exposure to domain-specific examples, leading to low accuracy paired with high confidence in clinical NLI tasks. Prompting techniques that improved general performance fail to reduce overconfidence in specialized domains.
Two-phase inference model shows knowledge retrieval operates in lower network layers while reasoning adjustment happens in higher layers. This separation explains why reasoning training improves math but can degrade knowledge-intensive domains like medicine.
RL enhances medical reasoning by suppressing incorrect domain knowledge during reasoning—not by expanding what models know. Evidence shows RL achieves +12.4 point knowledge improvement by removing low-reward reasoning trajectories that invoke wrong facts.
Show all 9 sources
RLAG rewards both answer accuracy and explanation rationality by cycling between augmented and unaugmented generation, progressively internalizing coherent knowledge structures. This outperforms SFT because it prioritizes reasoning quality over token-level correctness.
Proxy-tuning closes 88-91% of the alignment gap while surpassing direct fine-tuning on knowledge tasks by leaving base model weights untouched. Direct fine-tuning corrupts knowledge storage in lower layers, whereas proxy-tuning applies distributional shifts that primarily affect reasoning and style.
Supervised fine-tuning improves final-answer accuracy on benchmarks but cuts Information Gain by 38.9 percent, meaning models generate correct answers through post-hoc rationalization rather than genuine inferential steps. Standard metrics miss this degradation because they only measure final correctness.
Research shows every adaptation method—from parameter-efficient tuning to knowledge graph curricula—has optimal conditions tied to specific domains. The key finding: visible benefits like performance gains often come with hidden degradation in reasoning faithfulness, capability transfer, and format flexibility.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains
- Eliciting Reasoning in Language Models with Cognitive Tools
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
- Embedding Domain Knowledge for Large Language Models via Reinforcement Learning from Augmented Generation
- Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?
- A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
- Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive Survey
- An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models