INQUIRING LINE

Can an AI diagnostician spend about 70% less on tests and still diagnose as well, as one study of 304 hard cases suggests?

What role does cost estimation play in steering diagnostic test ordering?

This explores whether keeping track of what tests cost changes which tests an AI diagnostic system orders, and whether that tradeoff hurts accuracy.


This explores whether keeping track of what tests cost changes which tests an AI diagnostic system orders, and whether that hurts getting the diagnosis right. The short answer is that the corpus has one strong data point and a lot of related ideas from a different field. The data point is striking. On 304 hard New England Journal of Medicine case records, putting an orchestration layer around the o3 model cut average diagnostic spending from about $7,850 to $2,397 per case. Accuracy actually rose slightly, from 78.6% to 79.9% Can orchestration strategies boost diagnostic AI without better models?. The model alone was not cheap; it ordered tests freely. Most of the savings came from the scaffold, the structure built around the model, and the gains carried over to other model families. That suggests cost discipline is something you design into the process. It isn't something a smarter model simply learns.

Why would spending less not cost accuracy? A useful way in comes from research on how AI models spend their own thinking time. Studies of inference compute find that a fixed budget is wasteful. Giving easy problems less effort and hard ones more beats simply using a bigger model Can we allocate inference compute based on prompt difficulty?. Other work finds that the specific search method matters less than the total budget and how well the system can judge which paths are worth pursuing Does the choice of reasoning framework actually matter for test-time performance?. Test ordering has the same shape. Each test is a step that buys information, and a good diagnostician asks whether that information is worth its price given what is already known. In that sense, estimating cost acts as a value judgment: it forces each step to justify itself.

There is a tension worth knowing about. LLMs help clinicians partly by widening their list of possible diagnoses. In one study of 302 NEJM cases, clinicians with LLM access got the right answer into their top 10 list 51.7% of the time, versus 36.1% with search alone, and the authors credit the model's broader scope Does LLM assistance help clinicians build better differentials?. A wider list is good for catching rare diagnoses, but every added possibility invites another test. Cost awareness is the natural counterweight: it lets a system think broadly while testing selectively. That may be part of why a raw strong model spends so much and an orchestrated one spends far less.

Where the corpus runs thin: it doesn't show how the cost estimates are made, how accurate they are, or whether cost-steered ordering ever skips a test that mattered. Watch for that last risk. When AI suggestions are wrong, people defer to them heavily. In one study, wrong labels shown as AI output pulled experienced radiologists' accuracy from 82% down to 45.5% How much does wrong AI advice harm radiologist accuracy?. A cost-conscious AI that confidently says a test isn't needed could carry the same weight with clinicians. The savings look real on benchmark cases. Whether they hold up safely in actual wards is a question this collection doesn't yet answer.


Sources 5 notes

Can orchestration strategies boost diagnostic AI without better models?

On 304 NEJM cases, MAI-DxO orchestration achieved 79.9% accuracy at $2,397 per case versus 78.6% at $7,850 for o3 alone. The gains transferred across model families, suggesting the benefit comes from the scaffold, not model weights.

Can we allocate inference compute based on prompt difficulty?

Research shows inference effectiveness varies dramatically by prompt difficulty. Reallocating the same total compute adaptively—giving easy prompts less and hard ones more—substantially outperforms larger models under uniform budgets.

Does the choice of reasoning framework actually matter for test-time performance?

Information-theoretic analysis shows BoN and MCTS converge in reasoning accuracy when controlling for total compute. Snowball errors accumulate per step regardless of framework; mitigation depends on search scope and reward function reliability, not the specific algorithm.

Does LLM assistance help clinicians build better differentials?

In a study of 20 clinicians on 302 NEJM cases, those with LLM access achieved 51.7% top-10 accuracy versus 36.1% without it. The authors attribute the gain to the LLM's wider differential scope, making lists more comprehensive.

How much does wrong AI advice harm radiologist accuracy?

A 27-radiologist study found that incorrect BI-RADS suggestions caused experienced radiologists to drop from 82% to 45.5% accuracy, while inexperienced readers fell from nearly 80% to below 20%, demonstrating automation bias in mammography screening.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.