Why can't a medical AI learn medicine once and answer from memory, and why does it need live access to current sources?
Why does medical knowledge require continuous access to current sources?
This explores why medical AI can't simply learn medicine once during training and then answer from memory. The question is why it seems to need ongoing access to outside, up-to-date sources.
This explores why medical AI can't learn medicine once and answer from memory forever, and why it seems to need ongoing access to outside sources. One caveat first: the collection doesn't directly study how fast medical facts go out of date. What it does explain is why medicine leans so heavily on facts in the first place, and why stuffing more facts into a model's weights is a poor way to keep up. When researchers split model performance into 'knowing the right facts' and 'reasoning well,' accuracy on medical tasks tracks knowledge much more closely. Math shows the opposite pattern (Does medical AI need knowledge or reasoning more?). In medicine, the answer usually depends on having the right fact, not on clever reasoning.
That has a surprising side effect. Knowledge and reasoning seem to sit in different parts of the network. Fact retrieval happens in the lower layers and reasoning adjustments in the higher ones. This helps explain why training a model to reason harder can make it better at math but worse at medicine (Why does reasoning training help math but hurt medical tasks?). So the instinct to make a medical model 'think more' can backfire. What it often needs is to know more, and to know it correctly.
Why not just keep fine-tuning new findings into the model? A formal result shows that the number of facts a model can memorize in its weights is capped by its size. Each new fine-tuning round also tends to overwrite older knowledge and erode general ability. Tool use, meaning looking things up, has no such cap: a small model can recall an unlimited number of facts through a simple lookup mechanism (Can models store unlimited facts without growing larger?). In a field where guidelines, drug data and evidence keep changing, the case for constant access to sources is partly a capacity argument, not only a freshness one.
Access alone isn't the whole answer, though. Models do better when they learn, step by step, when to retrieve and when to trust what they already know. Unneeded retrieval adds noise (When should language models retrieve external knowledge versus use internal knowledge?). Retrieval can also fail quietly: when a question doesn't share words with the fact it depends on, retrieval systems surface that fact far less often than a model that already has it in context (Why do retrieval systems fail on queries that never mention needed facts?). A clinical question phrased around symptoms may never match a source written around a diagnosis. One safeguard from another field is to answer only when the retrieved evidence supports the answer and to refuse otherwise. This gives up some coverage in exchange for not inventing things (Can RAG systems refuse to answer without reliable evidence?).
The less obvious takeaway is that current sources also keep people honest, not just models. In a study on real NEJM cases, clinicians with LLM help produced broader, more accurate lists of possible diagnoses than clinicians using search alone (Does LLM assistance help clinicians build better differentials?). But a fluent model can also become a loop that echoes back what you already believe. Powerful models increase the need to anchor answers in real data rather than reducing it (Do foundation models actually reduce our need for real data?). Seen this way, ongoing access to sources is less about patching a model's memory than about keeping AI-produced medical knowledge tied to evidence someone can actually check.
Sources 8 notes
The KI/InfoGain framework reveals that medical domain accuracy correlates more strongly with knowledge correctness than reasoning quality, while mathematical domains show the inverse pattern. This distinction has direct implications for which training strategies to prioritize in each domain.
Two-phase inference model shows knowledge retrieval operates in lower network layers while reasoning adjustment happens in higher layers. This separation explains why reasoning training improves math but can degrade knowledge-intensive domains like medicine.
A formal proof and experiments show in-weight memorization is bounded by model size, while tool-use enables unbounded factual recall through a simple circuit. In-weight finetuning also degrades general capability by overwriting prior knowledge.
DeepRAG models each reasoning step as a Markov Decision Process where the model learns when to retrieve versus rely on parametric knowledge. The 21.99% improvement comes from better-targeted retrieval and elimination of noise from unnecessary external knowledge.
InMind's benchmark shows that with memory in context, models answer 84% of indirect queries correctly, but retrieval systems answer only 14.4% when the same memory must be retrieved. The systems store and recall the facts on demand, yet fail to surface them when user requests never mention connecting terms.
Show all 8 sources
A multilingual RAG system for noisy historical newspapers succeeds by aggressively expanding retrieval while constraining generation to only grounded answers. The grounded-refusal prompt prevents hallucination when OCR errors and language drift degrade source quality, trading coverage for integrity.
In a study of 20 clinicians on 302 NEJM cases, those with LLM access achieved 51.7% top-10 accuracy versus 36.1% without it. The authors attribute the gain to the LLM's wider differential scope, making lists more comprehensive.
Powerful foundation models don't eliminate the need for real data—they heighten it. Without empirical anchoring, iterative prompt refinement creates epistemic circularity where users confirm their own beliefs rather than test them.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs
- Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains
- Continual Learning Mechanisms Compose for Long-Horizon Memorization
- Capabilities of Gemini Models in Medicine
- Sequential Diagnosis with Language Models
- Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?
- Provable Benefits of In-Tool Learning for Large Language Models
- DeepRAG: Thinking to Retrieval Step by Step for Large Language Models