INQUIRING LINE

Do an AI's scores on test cases predict how it performs in a real clinic, alongside real doctors and patients?

Can offline LLM evaluation predict performance in live clinical workflows?

This explores whether a model's scores on test cases, vignettes and benchmarks tell you how it will perform once it is working inside a real clinic, alongside real clinicians and patients.


This explores whether a model's offline scores (on test cases, vignettes and benchmarks) tell you how it will do inside a real clinic. The corpus has no study that tests the same model offline and then in live care to compare the two. What it does show is that offline results and live results measure different things. Offline tests mostly measure what the model can do. Live results depend just as much on how the model fits into the work around it.

The offline evidence looks strong. In vignette experiments and an emergency room study, one LLM outperformed hundreds of physicians on diagnosis, triage and management Can language models reason better than physicians at diagnosis?. On 302 hard NEJM cases, clinicians with LLM access reached 51.7% top-10 differential accuracy, against 36.1% for clinicians using search alone Does LLM assistance help clinicians build better differentials?. When a system did meet real patients, the picture got more mixed. AMIE took histories from 100 urgent-care patients without a single safety stop. But its management plans trailed physicians' plans on practicality and cost Can conversational AI safely take patient histories without supervision?. Most case-based tests don't score those qualities, because they only come up when a plan has to work for a particular patient, clinic and budget.

The most telling live result comes from Nairobi primary care. Across 39,849 visits, clinicians with an LLM safety net made 16% fewer diagnostic errors and 13% fewer treatment errors Can AI safety nets reduce errors in live clinical practice?. The authors credit the asynchronous design, the interface and active rollout, not model capability alone. Forecasting research outside medicine reaches a similar conclusion: the way the workflow is structured mattered more than raw model strength Can LLMs actually forecast time series better than we think?. If that holds in clinics, a benchmark score tells you little about the part that decides results. The same model could help in one deployment and do nothing in another.

There is a second gap: time. DELEGATE-52 found that models with similar scores on short tasks diverged sharply by the 25th round of a long delegated workflow Do short benchmarks predict how models perform over long workflows?. Even frontier models corrupted about 25% of document content over long relays, and the damage never levelled off Do frontier LLMs silently corrupt documents in long workflows?. The type of failure also changes with capability. Weaker models visibly delete content. Stronger models quietly corrupt it while the document still looks intact Does model capability change how documents degrade?. These findings come from document workflows, not clinical ones, but clinical records are exactly the kind of long, repeatedly handed-off document involved. A model that scores better offline might fail in ways that are harder to catch, not less often.

Offline evaluation is still useful for what it measures. LLM-generated ratings of therapy engagement held up against real outcomes such as symptoms and effort Can local language models rate therapy engagement reliably?, which shows benchmark-style validation can link to the real world. So offline tests are a reasonable first filter for capability. They can't tell you how a model will perform once workflow design, long sessions and quiet errors come into play. Only deployment studies of the Nairobi kind measure that.


Sources 9 notes

Can language models reason better than physicians at diagnosis?

In physician-adjudicated vignette experiments and an emergency room study, a large language model outperformed hundreds of physicians on differential diagnosis, reasoning, triage, and clinical management tasks across multiple touchpoints.

Does LLM assistance help clinicians build better differentials?

In a study of 20 clinicians on 302 NEJM cases, those with LLM access achieved 51.7% top-10 accuracy versus 36.1% without it. The authors attribute the gain to the LLM's wider differential scope, making lists more comprehensive.

Can conversational AI safely take patient histories without supervision?

A single-arm study found that AMIE, a conversational AI system, conducted real clinical histories from 100 patients without requiring a single safety intervention by human supervisors. Patient attitudes toward AI improved after the interaction, though management plans trailed physicians on practicality and cost.

Can AI safety nets reduce errors in live clinical practice?

In 39,849 clinic visits, clinicians with access to an LLM safety net made 16% fewer diagnostic errors and 13% fewer treatment errors than those without. The authors attribute these gains to asynchronous, interface-optimized design and active deployment strategies, not model capability alone.

Can LLMs actually forecast time series better than we think?

LLMs have stronger intrinsic forecasting ability than recognized, but only when workflows separate numerical reasoning from contextual reasoning. Monolithic prompting obscures this capability; structured decomposition surfaces it.

Show all 9 sources
Do short benchmarks predict how models perform over long workflows?

DELEGATE-52 evaluated models across 50-round-trip relays and found short-interaction performance does not predict sustained delegation accuracy. Models ranking similarly on single-turn tasks diverged dramatically by relay 25, revealing degradation curves invisible to standard benchmarks.

Do frontier LLMs silently corrupt documents in long workflows?

Even the strongest models (Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4) degrade documents by ~25% over long relay workflows across 52 domains. Degradation decelerates but never plateaus, and errors compound silently, remaining undetected in spot-checked outputs.

Does model capability change how documents degrade?

DELEGATE-52 shows weaker LLMs degrade documents through visible deletion, while frontier models degrade through subtle corruption that preserves surface integrity. This shift makes frontier failures harder to detect and potentially more dangerous at workflow scale.

Can local language models rate therapy engagement reliably?

LLEAP achieved reliability (omega=0.953) and valid correlations with motivation, effort, and symptom outcomes using Llama 3.1 8B to rate 1,131 therapy sessions, while keeping data locally stored.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.