INQUIRING LINE

Medical AI models often seem to beat general ones, but how much of that gap comes from how they're tested?

How much does prompt selection bias favor medical models over base models?

This explores how much of the reported advantage of medically specialized AI models over the general-purpose models they were built from comes from how they were tested, especially from using one fixed prompt that happens to suit the medical model better.


This explores how much of the reported advantage of medical LLMs over their general-purpose base models comes from evaluation choices, especially the prompt used for testing. The corpus has one direct answer, and the effect is large. When seven medical LLMs and two medical vision-language models were compared with their base models in the usual way, the medical models looked significantly better on 70.5% and 62.5% of tasks. Two changes shrank that to 9.4% and 6.3%: optimizing the prompt separately for each model, and reporting confidence intervals so that noise wasn't counted as a win Does medical pretraining actually improve model performance?. One caveat: the study applied both changes together, so these numbers don't tell us how much is due to prompts alone and how much to statistical honesty. Combined, though, they remove most of the apparent benefit of medical pretraining.

Why would prompt choice favor one model so much? Prompts don't work the same way across models. A benchmark of 23 prompts across 12 LLMs found that a technique that helps one model can hurt another. Step-by-step reasoning, for example, lowered accuracy in stronger models while rephrasing helped weaker ones Do prompt techniques work the same across all LLM tiers?. So if a study picks one prompt, often one written with the medical model in mind, it may be testing how well each model fits that prompt rather than how much medicine each model knows. Separate work suggests prompt sensitivity tracks model confidence: models that are unsure swing widely when the wording changes Does model confidence predict robustness to prompt changes?. A base model that is less sure about clinical wording would be hit hardest by a prompt that doesn't suit it.

What the reader may not expect is what this says about where the knowledge lives. Prompt optimization can't add knowledge a model lacks. It can only bring out what is already there Can prompt optimization teach models knowledge they lack?. So when a well-prompted base model matches its medical version, the medical knowledge was mostly already in the base model. Medical fine-tuning often teaches the model how to respond to clinical-style prompts rather than giving it new medical knowledge. This fits evidence that reasoning ability comes from broad procedural knowledge spread across diverse pretraining data, not narrow domain-specific documents Does procedural knowledge drive reasoning more than factual retrieval?. It also fits reports of a general-purpose model outperforming hundreds of physicians on diagnosis without medical specialization Can language models reason better than physicians at diagnosis?.

This doesn't mean general models are safe in specialized areas. In clinical inference tasks, models can be wrong while sounding very confident, and prompting tricks that improve general accuracy don't fix that overconfidence Why do language models fail confidently in specialized domains?. A fairer comparison removes a false advantage, but it doesn't make either model reliable. The practical lesson: when a paper reports that a domain-specialized model beats its base model, check whether each model got its own best prompt and whether the gains go beyond the error bars. Without both, most of the reported advantage may disappear.


Sources 7 notes

Does medical pretraining actually improve model performance?

Seven medical LLMs and two medical VLMs showed significant improvements over base models in only 9.4% and 6.3% of tasks respectively when prompts were optimized per model and confidence intervals were reported, down from 70.5% and 62.5% without these controls.

Do prompt techniques work the same across all LLM tiers?

A 23-prompt benchmark across 12 LLMs shows rephrasing and background-knowledge prompts boost cheap models, while step-by-step reasoning reduces accuracy in high-performance models. Task structure, not generic best practices, determines which prompts help.

Does model confidence predict robustness to prompt changes?

ProSA found that when models are highly confident, they resist prompt rephrasing; low confidence causes major output swings. Larger models, few-shot examples, and objective tasks all correlate with higher confidence and greater robustness.

Can prompt optimization teach models knowledge they lack?

Prompting works entirely within a model's pre-existing training distribution and cannot supply domain knowledge absent from training data. This creates a hard ceiling: no prompt strategy can compensate for missing foundational knowledge, only reorganize what already exists.

Does procedural knowledge drive reasoning more than factual retrieval?

Analysis of 5 million pretraining documents shows reasoning relies on broad, transferable procedural knowledge from diverse sources, unlike factual recall which depends on narrow, document-specific memorization of target facts.

Show all 7 sources
Can language models reason better than physicians at diagnosis?

In physician-adjudicated vignette experiments and an emergency room study, a large language model outperformed hundreds of physicians on differential diagnosis, reasoning, triage, and clinical management tasks across multiple touchpoints.

Why do language models fail confidently in specialized domains?

LLMs trained on general text lack sufficient exposure to domain-specific examples, leading to low accuracy paired with high confidence in clinical NLI tasks. Prompting techniques that improved general performance fail to reduce overconfidence in specialized domains.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.