Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?

Paper · arXiv 2411.04118 · Published November 6, 2024
Domain Specialization in LLMs

Several recent works seek to develop foundation models specifically for medical applications, adapting general-purpose large language models (LLMs) and vision-language models (VLMs) via continued pretraining on publicly available biomedical corpora. These works typically claim that such domain-adaptive pretraining (DAPT) improves performance on downstream medical tasks, such as answering medical licensing exam questions. In this paper, we compare seven public “medical” LLMs and two VLMs against their corresponding base models, arriving at a different conclusion: all medical VLMs and nearly all medical LLMs fail to consistently improve over their base models in the zero-/few-shot prompting regime for medical question-answering (QA) tasks. For instance, across the tasks and model pairs we consider in the 3-shot setting, medical LLMs only outperform their base models in 12.1% of cases, reach a (statistical) tie in 49.8% of cases, and are significantly worse than their base models in the remaining 38.2% of cases. Our conclusions are based on (i) comparing each medical model head-to-head, directly against the corresponding base model; (ii) optimizing the prompts for each model separately; and (iii) accounting for statistical uncertainty in comparisons. While these basic practices are not consistently adopted in the literature, our ablations show that they substantially impact conclusions. Our findings suggest that state-of-the-art generaldomain models may already exhibit strong medical knowledge and reasoning capabilities, and offer recommendations to strengthen the conclusions of future studies.

Introduction. Recent advances in autoregressive large language models (LLMs) and vision-language models (VLMs) have attracted interest from practitioners in medicine, where these models hold great potential to transform various aspects of clinical practice (e.g., medical diagnosis, information retrieval from clinical documents, patient triaging) (Fries et al., 2022a; Moor et al., 2023a). State-of-the-art performance on various medical benchmarks is typically achieved by massive-scale closed-source models, such as GPT-4 (OpenAI, 2023a,b), MED-GEMINI (Saab et al., 2024; Yang et al., 2024), and MED- PALM (Singhal et al., 2023a,b; Tu et al., 2024), often performing on par with humans on medical licensing exams and open-ended consumer health question-answering (QA) tasks. However, the general lack of transparency in these models, high API usage costs, and patient data privacy concerns make their integration into routine clinical workflows challenging (Marks and Haupt, 2023). To address such concerns, recent works have proposed cheaper, open-source alternatives through domain-adaptive pretraining (DAPT; Gururangan et al., 2020), where a pretrained open-source general-domain model—such as LLAMA (Touvron et al., 2023a,b; Meta, 2024) or MISTRAL (Jiang et al., 2023) in the language space, and LLAVA (Liu et al., 2023) or OPEN-FLAMINGO (Awadalla et al., 2023) in the vision-language space—is continually pretrained on biomedical (image-)text corpora from public sources such as PubMed and medical textbooks. While some prior works show that medical models pretrained from scratch only using domain-specific corpora can outperform those trained via DAPT, both in the context of BERTstyle encoder-only models (Devlin et al., 2019; Gu et al., 2021; Yang et al., 2022) and decoder models (Taylor et al., 2022; Luo et al., 2022; Hernandez et al., 2023; Bolton et al., 2024), the DAPT approach has become common practice, resulting in a trend where the release of a more capable generaldomain model is typically followed by the release of its medical counterpart.

Despite the widespread adoption of medical DAPT, the claimed improvements in performance are worth scrutinizing. While the story is intuitive, more recent base models (e.g., LLAMA-3- 8B (Meta, 2024)) already exhibit strong off-theshelf performance on medical benchmarks without any adaptation (e.g., Open Medical LLM Leaderboard (Pal et al., 2024)), and given a lack of transparency about the pretraining corpora used to train the general-domain model in the first place, they may already be trained on relevant medical text.

Perhaps more concerning is the lack of applesto-apples comparisons in the literature. First, medical models resulting from DAPT are often only compared against baselines with different architectures (e.g., CLINICAL-CAMEL-70B (Toma et al., 2023) vs. GPT-4 (OpenAI, 2023a)) and under inconsistent evaluation setups (e.g., MEDITRON- 70B (Chen et al., 2023) fine-tuned on MedQA (Jin et al., 2020) vs. non-fine-tuned MED42-V1-70B (Christophe et al., 2024)), which can confound the interpretation of results. Second, the common practice of using a single, fixed prompting setup (e.g., prompt format, choice of few-shot examples) for all models under evaluation also warrants concern, as LLM/VLM behavior is extremely sensitive to such design decisions (Jiang et al., 2020; Zhao et al., 2021; Ceballos-Arroyo et al., 2024), and the “optimal” choice of such details rarely correlates between different models (Sclar et al., 2024).

In this paper, we perform an apples-to-apples comparison that addresses these concerns, comparing seven medical LLMs and two medical VLMs against their general-domain base models. We find that, for all but one LLM pair—BIOMISTRAL-7B (Labrak et al., 2024) vs. MISTRAL-7B-INSTRUCT- V0.1 (Jiang et al., 2023), a pair of models that performs fairly poorly in absolute terms—the open-source medical LLMs and VLMs that we evaluate do not consistently improve over their general-domain counterparts on various medical (visual) QA tasks (Figure 1). We compare several pairs of general-domain and medically adapted LLMs/VLMs (see Table 1), whose only differences lie in medical DAPT (i.e., one model is the base model, from which the other is derived via medical DAPT). For each pair, we compare their performances from zero-/few-shot prompting (Radford et al., 2019; Brown et al., 2020), after independently selecting the “best” prompt format and few-shot examples for each model based on the validation set and accounting for statistical uncertainty in model comparison. Our findings (Section 4) suggest that state-ofthe-art general-domain models may already exhibit strong medical knowledge and reasoning capabilities that can be leveraged effectively when prompted appropriately. Our main contributions can be summarized as follows:

Related work. DAPT (Gururangan et al., 2020) is a transfer learning approach, where a pretrained model is further pretrained on domain-specific data for better alignment to a target domain of interest (e.g., medicine, law). Several studies show that language models trained via DAPT often outperform their generaldomain counterparts on domain-specific tasks, such as claim detection from blog posts (Chakrabarty et al., 2019), named entity recognition from German novels (Konle and Jannidis, 2020), and judgment prediction for legal cases (Xiao et al., 2021). In the medical domain, prior works based on BERTstyle encoder-only language models (Devlin et al., 2019), such as BIOBERT (Lee et al., 2019) and CLINICALBERT (Alsentzer et al., 2019), show that medical DAPT improves fine-tuning performance on tasks such as medical concept extraction from patient reports (Uzuner et al., 2011), identification of gene-disease relations from PubMed abstracts (Do ̆gan et al., 2014; Bravo et al., 2015; Krallinger et al., 2017), and natural language inference on clinical notes (Romanov and Shivade, 2018). More recent works suggest that decoder-based autoregressive LLMs and VLMs trained via medical DAPT also show strong performance on various medical tasks. Medical LLMs such as MED- ITRON (Chen et al., 2023), adapted from LLAMA- 2 (Touvron et al., 2023b); and BIOMISTRAL (Labrak et al., 2024), adapted from MISTRAL- 7B-INSTRUCT-V0.1 (Jiang et al., 2023); perform well on knowledge-intensive QA tasks based on medical licensing and academic exams (Jin et al., 2020; Pal et al., 2022; Hendrycks et al., 2021) and PubMed abstracts (Jin et al., 2019). Medical VLMs such as LLAVA-MED (Li et al., 2023), adapted from LLAVA (Liu et al., 2023); and MED- FLAMINGO (Moor et al., 2023b), adapted from OPEN-FLAMINGO (Awadalla et al., 2023); also perform well on visual QA tasks based on radiology (Lau et al., 2018; Liu et al., 2021) and pathology images (He et al., 2020) and academic exams (Yue et al., 2024). These encouraging results have established DAPT as a go-to approach for training a medically specialized model, a conclusion that we re-examine in this work.

Method. 3 Experimental Setup Evaluation Metric. Since we focus on closedended QA tasks, we use exact-match accuracy as our main evaluation metric. Following the Holistic Evaluation of Language Models (HELM) benchmark (Liang et al., 2023), when we consider greedy decoding, we treat the text generated by a model (without any constraints on the vocabulary) to be its prediction, and check for an exact match between the prediction and the correct answer up to primitive string operations (e.g., lower-casing, removing white space/punctuation). To handle cases where the model simply repeats the list of answer choices or produces an ambiguous answer (e.g., selecting multiple answer choices), we take a conservative approach and treat the prediction to be incorrect, even if there is a match. Meanwhile, to quantify the extent of improvement from medical DAPT, we also consider the relative accuracy of the medical model with respect to the general-domain model. Formally, we define relative exact-match accuracy as E[1[fmedical(x) = y] −1[fgeneral(x) = y]] ∈ [−1, 1], where fmedical and fgeneral denote the medical and general-domain models, x and y denote the input prompt and answer in a QA pair from the test set, and 1[·] denotes the indicator function. This metric quantifies the difference in accuracy between the medical model and the general-domain model. To distinguish the two metrics, we refer to the former as the absolute exact-match accuracy in subsequent discussions.

Assessing Statistical Significance. Given the relatively small size of test datasets in medical QA benchmarks, it is important to assess whether the In this section, we provide an overview of our approach to assess whether medical DAPT leads to statistically significant improvements in zero- /few-shot medical QA performance. For few-shot prompting, we consider the 3-shot setting to ensure that the input prompt is shorter than the context window sizes for all models evaluated. For evaluation, we pay special attention to two aspects. First, language models are highly sensitive to the choice of prompting strategy (e.g., prompt format, choice of few-shot examples), where seemingly insignificant changes to the prompt can lead to idiosyncratic model behavior (Jiang et al., 2020; Zhao et al., 2021). Second, prior works show that the “optimal” choice of prompt format rarely correlates between different models (Sclar et al., 2024), suggesting that using a single, fixed prompt for all models for comparison can result in misleading conclusions. To ensure a fair comparison that isolates the impact of medical DAPT, we treat the choice of prompt format and few-shot examples as additional hyperparameters when generating predictions, and

Discussion. Here, we summarize the main findings from the zero-/few-shot prompting experiments outlined in Section 3. Unless specified otherwise, we focus on the greedy decoding results in subsequent discussions and include the results for constrained decoding in Appendix E. Overall, we find that all medical VLMs and nearly all medical LLMs fail to consistently improve over their general-domain counterparts in the zero-shot and few-shot prompting regimes. Moreover, we demonstrate the importance of rigorous experimental design in surfacing this finding—performing pairwise model comparison with a single, fixed prompt optimized only for the medical model, while ignoring statistical uncertainty, paints a misleadingly optimistic picture of medical DAPT performance.

Finding 1: After model-specific prompt selection, the vast majority of medical models fail to consistently show a statistically significant improvement over the general-domain models. In Figures 3–4, we show the absolute and relative exact-match accuracies achieved by the medical and general-domain LLMs and VLMs in the zero- /few-shot prompting regime. For LLMs, we only show the 3-shot prompting results in the main text (see Appendix D for results in the zero-shot setting, which are similar). We exclude the results for CLINICAL-CAMEL-70B on both versions of MedQA, as the model has already been trained on a subset of the official training split (see Table 1 in Toma et al. (2023)). For VLMs, we show both zero-shot and 3-shot results, as LLAVA-V0-7B and LLAVA-MED-7B were not pretrained to handle inputs with multiple images. We calculate the confidence intervals via bootstrapping on the test set, as described in Section 3.

Finding 2: Using a single, fixed prompt for all models and overlooking statistical uncertainty may overestimate the performance benefits of medical DAPT. Based on Finding 1, we further 5 Discussion and Conclusion In this work, we investigated the effectiveness of DAPT for training medically specialized LLMs and autoregressive VLMs suitable for knowledgeintensive medical (visual) QA tasks. To that end, we compared several pairs of state-of-the-art medical LLMs/VLMs to their general-domain counterparts, whose only differences lie in medical DAPT and are exactly identical in model architecture and scale. Our work diverges from prior works by providing a direct apples-to-apples comparison of medical and general-domain models while accounting for LLM/VLM sensitivity to prompting details and assessing the statistical significance of the results. Across both model classes and all model scales, we found that the performance benefits from medical DAPT largely disappear when we (i) tailor the prompt format and choice of few-shot examples to each medical and general-domain model separately; and (ii) account for statistical uncertainty in model comparison. In particular, we found that when we optimize the prompt only for the medical model and compare each model pair based on their absolute accuracies without accounting for uncertainty, the performance improvements from medical DAPT can be overestimated, potentially leading to unreliable conclusions about the benefits of medical DAPT. For example, in the zero-shot setting, evaluation under this setup leads to the conclusion that medical LLMs and VLMs, on average, outperform the corresponding general-domain models in 70.5% and 62.5% of all QA tasks, while the improvements are in reality statistically significant in only 9.4% and 6.3% of tasks after optimizing the prompt for each model to ensure a fair comparison. Our findings suggest that for state-of-the-art general-domain LLMs and VLMs, the performance benefits from additionally pretraining on medical data from public sources such as PubMed may be limited.

Limitations. We discuss our findings with the following caveats. First, there is a vast and growing set of papers on applying medical DAPT to various general-domain base models, and we could not hope to compare all publicly available models here. While we selected the models to cover a wide range of general-domain base models and model scales (7B–70B) (Table 1) and included some of the latest models (e.g., OPEN- BIOLLM and LLAMA-3), it is always possible that some newly released models do in fact yield better zero- or few-shot performance on medical QA. Second, we focus in this paper on the narrower task of closed-ended medical QA. In part, this choice reflects the fact that such benchmarks are well-standardized and highly publicized. However, they do not reflect the breadth of possible applications of LLMs and VLMs in medical domains. For instance, Singhal et al. (2023b) show that medical LLMs such as MED-PALM-2 can produce physician-level answers to open-ended consumer health queries, and Agrawal et al. (2022) demonstrate the potential of using LLMs for extracting information from structured clinical notes. Some would argue that such tasks are a more realistic application of such models in practice, and it is certainly possible that an analysis like ours would find improved performance on such tasks, though we do not investigate these tasks in the present work.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do curriculum design and feedback approaches affect model learning? How do clinicians calibrate trust in AI medical recommendations? How does fine-tuning trade off accuracy against reasoning quality? Can language models reliably simulate personas and predict behavior? Why do language models fail at sustained therapeutic relationships despite understanding techniques? What limits language model accuracy in evaluating ideas? Can confidence signals reliably detect flawed reasoning in language models? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? What prevents LLMs from applying their reasoning knowledge to improve outputs? Can external verification systems adequately replace learned reasoning in AI outputs? What explains the gap between benchmark scores and true reasoning capability? Can smaller specialized models match frontier models on key metrics? Can base models hide emergent misalignment through alignment training?