AI-based Clinical Decision Support for Primary Care: A Real-World Study
We evaluate the impact of large language model-based clinical decision support in live care. In partnership with Penda Health, a network of primary care clinics in Nairobi, Kenya, we studied AI Consult, a tool that serves as a safety net for clinicians by identifying potential documentation and clinical decision-making errors. AI Consult integrates into clinician workflows, activating only when needed and preserving clinician autonomy. We conducted a quality improvement study, comparing outcomes for 39,849 patient visits performed by clinicians with or without access to AI Consult across 15 clinics. Visits were rated by independent physicians to identify clinical errors. Clinicians with access to AI Consult made relatively fewer errors: 16% fewer diagnostic errors and 13% fewer treatment errors. In absolute terms, the introduction of AI Consult would avert diagnostic errors in 22,000 visits and treatment errors in 29,000 visits annually at Penda alone. In a survey of clinicians with AI Consult, all clinicians said that AI Consult improved the quality of care they delivered, with 75% saying the effect was “substantial”. These results required a clinical workflow-aligned AI Consult implementation and active deployment to encourage clinician uptake. We hope this study demonstrates the potential for LLM-based clinical decision support tools to reduce errors in real-world settings and provides a practical framework for advancing responsible adoption.1
Introduction. Artificial intelligence (AI) systems have the potential to widen access to reliable health information and highquality care (Beam and Kohane, 2018; Topol, 2019; Rajkomar et al., 2019). Large language models (LLMs) have recently experienced significant leaps in performance, reliability, and safety for health applications (Arora et al., 2025; Nori et al., 2025; Singhal et al., 2023, 2025). These advances suggest new opportunities for improving healthcare delivery—including supporting clinicians in delivering better care.
Despite research progress, scaled real-world deployment of AI tools in clinical environments remains limited. State-of-the-art LLMs now often outperform physicians on benchmarks (Goh et al., 2025; Arora et al., 2025; Nori et al., 2025; Van Veen et al., 2024), but these gains have yet to translate into measurable benefits for patients and clinicians in live care settings. The most critical bottleneck in the health AI ecosystem is no longer better models, but rather the model-implementation gap: the chasm between model capabilities and real-world implementation.
Closing the model-implementation gap necessitates the responsible study of LLM implementations in frontier health AI use cases. One example is clinical decision support (CDS) systems (Sutton et al., 2020; Middleton et al., 2016), which provide clinicians with relevant knowledge at the point of care. Efforts to measure how well LLMs can help with clinical decisions so far have used offline evaluations, often measuring model capabilities on clinical vignettes without capturing the unique challenges of designing and deploying an implementation for real-world care (Benary et al., 2023; Goh et al., 2025; Oniani et al., 2024).
In this study, we examine the impact of an LLM-based clinical decision support tool in live care. Penda Health, where several authors are affiliated, is a network of high-volume clinics in Nairobi, Kenya that delivers 24-hour primary and urgent care to a broad range of Nairobi residents. We studied Penda’s AI Consult, which serves as a clinical safety net to prevent errors. The system is triggered asynchronously during key clinical workflow decision points in the electronic medical record (e.g., diagnosis, treatment). It surfaces guidance through a tiered traffic-light interface (green: no action, yellow: advisory, red: requires review), and is explicitly designed to minimize cognitive burden and preserve clinician autonomy. The tool was developed through iterative co-design with frontline clinicians and tailored to local epidemiology, Kenyan clinical guidelines, and Penda’s care protocols.
To assess the tool’s impact, we conducted a pragmatic cluster-assigned study of 39,849 visits, comparing outcomes for patient visits managed by clinicians with and without access to AI Consult. We aimed to evaluate three primary domains: (i) clinical quality, as rated by independent physicians reviewing clinical documentation with patient identification removed; (ii) use and usability, based on a clinician survey and AI Consult usage data; and (iii) patient-reported outcomes collected via routine follow-up calls. We find meaningful reductions in clinical errors for clinicians with the tool (“AI group”) vs those without (“non- AI group”) and encouraging feedback from clinicians using AI Consult. We did not detect a significant difference in patient-reported outcomes. This study was conducted with the approval and consultation of Kenya’s Ministry of Health, Kenya’s Digital Health Agency, Nairobi County, AMREF Health Africa Ethical and Scientific Review Committee, Kenya’s National Commission for Science, Technology and Innovation (NACOSTI), and other local stakeholders to ensure it aligned with national priorities, ethical standards, and data protection requirements.
• We describe the key factors for success: a capable model, a clinically-aligned implementation, and active deployment strategies.
This work offers an early demonstration of the potential for LLM-based tools to serve as real-time copilots for delivering care and a practical framework for advancing responsible adoption in real-world health systems.
Related work. Offline evaluation of LLMs for health. Advances in LLMs have spurred many works evaluating them for health applications. Prior works have evaluated health performance broadly (Arora et al., 2025; Bedi et al., 2025) or for specific tasks, including differential diagnosis (McDuff et al., 2025; Nori et al., 2025; Goh et al., 2025), clinical summarization (Van Veen et al., 2024; Zaretsky et al., 2024), radiology report generation (Tanno et al., 2025; Tu et al., 2024), and Q&A (Ayers et al., 2023; Nori et al., 2023; Pfohl et al., 2024). Some works have focused on specialized models (Moor et al., 2023; Li et al., 2023; Singhal et al., 2023, 2025; McDuff et al., 2025; Tu et al., 2025, 2024; Saab et al., 2024; Yang et al., 2024) and others on general models (Ayers et al., 2023; Nori et al., 2025, 2023; Saab et al., 2025; Arora et al., 2025; Johnson et al., 2023). Evaluation in many works relies heavily on narrow automated benchmarks that measure clinical knowledge (Nori et al., 2023; Singhal et al., 2023). Some works have evaluated models across many benchmarks, offering more robust characterizations of model performance across tasks (Bedi et al., 2025; Saab et al., 2024; Tu et al., 2024). Other works have employed human evaluation with physicians or patients, sometimes employing realistic clinical vignettes or electronic medical record data (Goh et al., 2025; Ong et al., 2024; Dash et al., 2023; Ayers et al., 2023; Singhal et al., 2025; Pfohl et al., 2024). Some recent works have combined human and automated evaluation towards clinician-aligned evaluation at scale (Arora et al., 2025; Fleming et al., 2024). All of these works involve “offline” evaluation of LLMs, which do not enable the study of the unique challenges of bringing model advances into clinical practice, including real-world patient diversity, designing for and learning from clinician workflows, and deployment towards successful clinician uptake. Unlike prior evaluations of LLMs, the present study examines outcomes of using an LLM-based tool live during patient care at scale, addressing the unique challenges of real-world implementation.
Clinical decision support. AI Consult is an example of a clinical decision support system. Such systems have been used in various forms since the 1970s (Sutton et al., 2020; Middleton et al., 2016; Shortliffe, 1977; Bright et al., 2012; Musen et al., 2021). These systems support clinicians with knowledge and tools at the point of care.
Method. Primary care clinicians see patients across every age group, organ system, and disease type, often in the same day, requiring broad knowledge. The breadth of practice contributes to primary care quality challenges worldwide, with the WHO reporting substantial rates of preventable patient harm (WHO, 2023). This suggests that AI systems could be especially useful in primary care.
In Kenya, primary care is largely delivered by clinical officers: clinicians who complete three years of academic training followed by a one-year supervised internship. They manage the full breadth of acute and chronic conditions across the life course. Structural challenges in Kenyan primary care (late presentation, high patient volumes, limited diagnostics) compound this wide scope of practice to create a sizable quality gap: Studies suggest low adherence to national guidelines by healthcare workers across multiple levels of Kenya’s healthcare system, with frequent errors such as missed comorbidities, antibiotic overprescription, and diagnostic delays (Marete et al., 2020; Kr ̈uger et al., 2017; Kiener et al., 2025).
Penda Health is a Nairobi-based social enterprise founded in 2012 that delivers comprehensive, 24-hour primary and urgent care services through a network of fully-licensed medical centers distributed across the city. The organization presently operates 16 clinics and records over 1000 patient visits a day, supported by a clinical workforce of more than 100 licensed clinical officers. For a video and photos depicting Penda’s care context and AI Consult, see the blog post that accompanies this paper.
Penda has invested substantially in its digital infrastructure and quality improvement programs over the years, and has been a pioneer in implementing clinical decision support tools.
Electronic medical record (2017). A cloud-hosted electronic medical record (EMR), Easy Clinic, was introduced in 2017, supporting all patient visits and enabling real-time monitoring of quality metrics and operations.
Rule-based system (2019-2020). Penda implemented an early non-AI CDS system before its first iteration of AI Consult (Korom and Njue, 2020). In this system, decision trees embedded in the EMR provided point-of-care reminders for some common conditions. Similar early approaches have been employed and studied in other contexts for many years (Papadopoulos et al., 2022; Bright et al., 2012; Musen et al., 2021).
AI Consult v1 (February 2024). Penda Health implemented an early version of an LLM copilot prior to the version studied in this work. AI Consult v1 provided feedback from an LLM on the current visit at clinician request. Clinicians clicked a button within the EMR during a patient visit, chose an area to receive feedback on (including documentation, patient management, and overall visit), and received structured feedback from an LLM. Similar to Penda’s other CDS iterations, clinicians reviewed the output of the tool and made all clinical decisions.
During the early deployment of AI Consult v1, Penda performed an internal safety audit of 100 randomly selected cases. Each of these cases included (1) patient documentation state before AI Consult use, (2) AI Consult response, and (3) final documentation state, including any changes resulting from AI Consult. These cases were reviewed by Penda’s quality team. Each AI Consult output was scored from 1–5, where 5 was outstanding feedback from the LLM on the case (relevant, locally-appropriate, comprehensive, and actionable); 3 was neutral; and 1 was actively harmful (e.g., encouraging the clinician to perform unnecessary tests, offering an inappropriate diagnosis, or an incorrect or not locally appropriate treatment plan). Cases were also annotated with qualitative notes on how clinicians may have acted on AI Consult responses.
Despite showing early promise in terms of patient safety and quality improvement, AI Consult v1 only achieved adoption in about 60% of visits. Qualitative notes showed many cases in which AI feedback was not heeded despite being correct and clinically actionable.
Discussion. Our findings demonstrate that a large language model–based clinical decision support tool can meaningfully reduce diagnostic and treatment errors when deployed in live outpatient care. This improvement occurred not in simulation or review of EMR data, but in the context of routine, real-world practice across nearly 40,000 patient visits in 15 clinics—supporting our view that AI systems, when carefully implemented in clinician workflows, can enhance care quality.
The scale and scope of AI Consult are also notable. Unlike prior decision support systems which target narrow conditions, specialties, or workflows—such as drug interactions or chronic disease screening—AI Consult operated continuously, across all patient visits and key decision points.
One of the most important implications of this work is the potential for AI tools to further improve the quality of care delivered by primary care clinicians. By functioning as an asynchronous safety net and surfacing real-time feedback at decision points, AI Consult provides lightweight supervision that improved care without undermining clinician autonomy. In this sense, the system serves not only as a quality assurance mechanism but as an empowering tool for clinicians.
Beyond reducing errors in real-time, AI Consult appeared to foster substantial skill gains. During the study period, the proportion of visits that “started red”—a proxy for clinicians missing a critical issue on first pass—for treatments specifically fell by about 10–15 absolute percentage points in the AI group while remaining flat in the non-AI group. Because these initial alerts precede AI feedback on treatments, the decline signals that clinicians internalized the system’s feedback and preemptively avoided common failure modes. The magnitude of the effect is notable; for every 7-10 patients AI group clinicians saw, they avoided one important initial treatment error. Such learning effects were evident not only for treatment decisions but also for history-taking, suggesting AI Consult facilitates broader learning rather than narrow protocol adherence. These findings, together with survey responses citing the tool as “very informative,” “a learning tool,” and helpful in “sharpening my skills,” support the view that well-designed copilots can function as continuous, case-based education—uplifting individual competence while simultaneously safeguarding patients.
Clinically-aligned implementation was a key factor in the effectiveness of AI Consult. Penda’s previous iteration of AI Consult (Section 2.3) achieved limited uptake because it required clinicians to interrupt the flow of a patient visit to request AI feedback. The iteration we studied here provided a tiered, low-friction interface, enabling broad coverage with minimal disruption and alert fatigue. These changes reflect learnings from the implementation science literature, which has found that avoiding alert fatigue and surfacing CDS recommendations automatically instead of on demand improved clinician adherence (Kawamoto et al., 2005; Van de Velde et al., 2018; Seidling et al., 2011). Clinician feedback affirmed the utility of the tool–all AI group survey respondents reported that AI Consult improved the quality of care they could deliver–and indicated overall enthusiasm (“It should be provided to all health care providers”).
Active deployment was another key factor for the success of AI Consult. The tool had a significantly greater effect during the main study period (when active deployment strategies were employed) compared to the induction period (Table 4), with a clear divergence between AI and non-AI groups for left in red rate and started red rate over the seven weeks of the main period (Fig. 6).
Conclusion. We have presented a real-world evaluation of a large language model-based clinical decision support tool deployed in live patient care, with a meaningful reduction in diagnostic and treatment errors. Our findings underscore three critical components: (1) capable models, which are now widely available; (2) clinicallyaligned implementation, which supports the user rather than distracting them; and (3) active deployment, including building clinician connection, measurement, and incentives. Clinical impact does not emerge solely from model performance, but from a confluence of technical, human, and organizational factors.
With advancements in model capabilities, closing the model-implementation gap has become the most important challenge for the health AI ecosystem. This study provides a template for how AI systems can be safely and effectively embedded into clinical workflows. Further progress requires coordinated efforts across the ecosystem, including policymakers developing regulatory frameworks, engineers designing better implementations, and healthcare systems driving thoughtful deployments. Ultimately, we hope that systems like AI Consult will become the standard of care, supporting clinicians in delivering safer, more consistent, and more accessible care worldwide.
Limitations. 5.1 Limitations AI Consult represents an early, promising archetype of an AI-powered clinical copilot. While the results are encouraging, we emphasize that this is a first step. Continued iteration will be essential—to reduce documentation burden, improve contextual relevance, and align more closely with local practice norms. Future implementations may include voice-first interfaces, real-time charting assistants, or agents that execute clinician-confirmed actions in electronic medical records.
Although AI Consult was associated with reduced diagnostic and treatment errors, we did not observe statistically significant differences in patient-reported outcomes during the study period. This may reflect limitations in measurement sensitivity, response rates (response rates were 40%), short follow-up period, or the relatively short duration of the study. Further work—particularly large studies powered for patient outcomes—will be needed to assess the downstream impact of AI-assisted care.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What prevents LLMs from applying their reasoning knowledge to improve outputs?- Can offline LLM evaluation predict performance in live clinical workflows?
- Can LLM performance on zebra cases predict results in routine clinical practice?
- Does medical fine-tuning help LLMs through knowledge or reasoning ability?
- How do LLM performances compare across different types of medical tasks?
- How much does missing images and tables limit LLM diagnostic reasoning?
- Why do clinicians fail to act on correct AI suggestions in real care?
- What evidence would prove medical AI actually works in clinics?
- What proportion of patients cannot complete AI interviews due to technology barriers?
- Why did primary care physicians review only 73% of AI-generated transcripts?
- How much does self-play training with LLM-simulated patients actually improve diagnostic accuracy?
- Does AMIE's advantage hold when patients interact through speech or video instead of text?
- Can an AI system trained on text consultations handle diagnostic uncertainty in real patient encounters?
- Does this colonoscopy finding apply to other medical specialties using AI?
- How much do physician scores improve when assisted by the same model?
- Why did lay users with AI models fail to match unaided physicians on diagnosis?
- Can rubric-graded response quality predict real-world clinical workflow success?
- Do consensus criteria identify behaviors where physicians and models differ most?
- Does AI change clinician cognition or just increase reliance on predictions?
- Why do radiologists fail to benefit from AI decision support?
- Do patients actually perceive AI as worse at addressing their unique medical needs?
- What prospective trials are needed to validate AI diagnostic claims?
- How much diagnostic accuracy is gained when physicians receive expert advice?
- Why does medical knowledge require continuous access to current sources?