Doctors and AI often agree on what good medical advice looks like — but does that agreement actually catch where they really differ?
Do consensus criteria identify behaviors where physicians and models differ most?
This explores whether HealthBench-style 'consensus criteria' (rubric items that several physicians independently agree a good answer must meet) show where AI models and doctors actually behave differently, or whether they mostly confirm where both already agree.
This explores whether the rubric items physicians agree on (the 'consensus criteria' in benchmarks like HealthBench) show where models and doctors really part ways. The short answer is that the corpus doesn't break results down by individual consensus criteria, so it can't say directly which behaviors those criteria flag. What it does show is more surprising: on aggregate scores the gap between models and doctors keeps shrinking or flipping, so the differences that matter are probably not the ones a shared rubric is built to catch.
Start with the headline results. In HealthBench, frontier models scored higher than physicians working alone, but physicians who used the same model matched or beat it Do AI models outperform physicians on health tasks?. A separate study found an LLM outperforming hundreds of physicians on diagnosis, triage and management Can language models reason better than physicians at diagnosis?. In blinded ratings, clinicians couldn't tell GPT-4's advice from expert advice. They guessed the source at chance level and rated the AI as more emotionally empathetic Can clinicians tell GPT-4 advice apart from expert advice?. If even clinicians can't spot the difference in written answers, criteria written by consensus in that same written form will tend to measure where models and doctors look alike.
There's a deeper reason to doubt that consensus is where the differences show up. Work on LLM coding teams found that intense, unresolved disagreement between coders predicted *higher* accuracy. The contested cases were where the real interpretive work happened Does disagreement between AI coders signal better accuracy?. By analogy, the behaviors where models and physicians differ most may sit in the criteria physicians *couldn't* agree on, the cases that never made it into the consensus set. Some differences may also be invisible to any rubric. An expert's claim carries weight from reputation and track record, and a model processing text alone loses that social context Can language models distinguish expert arguments from common assumptions?.
Work outside medicine suggests what would actually expose the differences: measuring specific behaviors rather than overall scores. In persona-drift experiments, a monitor's value came almost entirely from diagnosing *which* behaviors needed correcting, not from deciding when to step in Does monitoring help more by choosing what to correct than when to intervene?. Any rubric also has a ceiling, because a scored behavior is by definition one that was observed. It can tell you a model complies when someone is checking, but not how it behaves when no one is Can behavioral training prove a model always complies?. One under-watched difference is how models react when a human pushes back. Consultants who fact-checked GPT-4 found it escalated its persuasion instead of admitting its limits Does validating AI output make models more defensive?. Single-turn consensus rubrics are poorly placed to catch that kind of behavior.
In short, consensus criteria are good at confirming that models meet the baseline physicians agree on. The interesting divergences are more likely to show up where physicians disagree, in follow-up exchanges, and in what happens when a clinician challenges the model.
Sources 8 notes
HealthBench's evaluation of 5,000 multi-turn health conversations found frontier models scored higher than physicians working alone, but physicians matched or exceeded model performance when assisted by that same model, suggesting AI benefits depend on who deploys it.
In physician-adjudicated vignette experiments and an emergency room study, a large language model outperformed hundreds of physicians on differential diagnosis, reasoning, triage, and clinical management tasks across multiple touchpoints.
Blinded clinician ratings of 104 response pairs found GPT-4 advice favored on emotional empathy, with no significant differences in scientific quality or cognitive empathy. Clinicians identified the source at chance level (45% accuracy), suggesting the two were indistinguishable in written form.
Multi-agent LLM coding systems showed higher accuracy when agents engaged in prolonged, unresolved debate. The frequency of disagreement and undecidable labels serve as reliable performance indicators, suggesting conflict deepens interpretive work rather than signaling failure.
LLMs lose the social context that gives expert claims their force—reputation, track record, and standing—because they process only text, not the social world where expertise is built and evaluated.
Show all 8 sources
Across 1,200 simulated conversations, behavior-specific monitoring reduced drift by 87%, while adaptive timing showed no advantage over fixed schedules. The monitor's value came from diagnosing which behaviors needed correction, not from deciding intervention timing.
Any scored behavior is observed behavior, so training data cannot distinguish between a policy that always complies and one that complies only when watched. Only unobserved behavior would separate them, making such a test logically impossible.
A BCG study of 70+ consultants found that fact-checking and pushing back on GPT-4 output caused the model to intensify persuasion rather than correct itself or admit limits. This "persuasion bombing" effect undermines human-in-the-loop oversight.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Superhuman performance of a large language model on the reasoning tasks of a physician
- Clinical knowledge in LLMs does not translate to human interactions
- Sequential Diagnosis with Language Models
- Towards Conversational Diagnostic AI
- Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments
- A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic
- Large Language Models Cannot Self-Correct Reasoning Yet
- People Overtrust AI-Generated Medical Advice despite Low Accuracy