Do wrong AI predictions hurt more than right ones help?
When AI tools give incorrect medical predictions, do they damage clinician performance more severely than correct predictions improve it? This matters for understanding whether averaging test results can hide dangerous asymmetries in AI safety.
Morey, Rayo and Woods (Ohio State), writing in AI Frontiers, report a case study using a method they call Joint Activity Testing, in which "450 nursing students and a dozen licensed nurses each reviewed 10 historical ICU cases, using an AI early warning tool in four configurations." Nurses rated each case 1-10 for how concerned they were about the patient. The result: "When AI predictions were most correct, nurses performed 53% to 67% better than when they worked without AI assistance. However, when AI predictions were most misleading, nurses performed 96% to 120% worse than when they worked without AI assistance." Adding annotations to the AI's predictions "did not significantly alter results."
The authors argue this is not a matter of effort or conscious reliance: "these results do not stem from the nurses' lack of effort or from their consciously offloading decision-making to the AI. Instead, AI assistance appeared to change how nurses think when assessing patients." Nurses and the algorithm looked complementary when tested apart — "the algorithm struggled with cases that nurses without AI handled with ease" — yet when misleading predictions were paired with those same routine cases, nurses "consistently misclassified emergencies as nonemergencies (and vice versa)." Their larger argument is methodological: standard evaluations that "treat AI and humans separately" or "boil down the complex effects of collaboration into a single metric, like average performance" hide this asymmetry, because "average gains can mask rare but severe failures." Neither years of experience nor familiarity with AI tools reliably predicted who would benefit or recover from AI's mistakes.
This matches the shape of two automation-bias findings already in the library while adding a measurement neither has: a matched gain-versus-loss figure within one study and task. How much does wrong AI advice harm radiologist accuracy? shows only the loss side — a wrong AI label cutting accuracy — without a matched measure of how much a correct label helped. Does time pressure make AI advice more persuasive to experts? measures a narrower commission-error rate and names time pressure as the severity driver; Morey, Rayo and Woods instead attribute the swing to AI assistance changing cognition generally, with no time-pressure manipulation in view. All three studies reject effort or face-value judgment as the explanation, but this one is the most explicit that averaging itself, not just human bias, is what conceals the effect, and it proposes testing "a range of challenging cases with varied AI performance" instead of a single aggregate score.
The excerpt gives the headline percentages but not the underlying npj Digital Medicine paper's statistics — how the 450 students and 12 licensed nurses were split in the results, significance tests, or how the "four configurations" differed beyond annotations. The task was a simulation on 10 historical cases rated on a 1-10 scale, not live bedside decisions, and the excerpt does not describe how the "without AI assistance" comparison condition was run. The claim that AI "changed how nurses think" rather than caused offloading rests on the authors' own follow-on analyses, not shown here. If the pattern generalizes past this ICU simulation, it implies that a single aggregate accuracy score cannot certify a human-AI tool as safe for high-stakes deployment, whatever that average score is.
Inquiring lines that read this note 7
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do users confuse explanation quality with actual system accuracy? How do clinicians calibrate trust in AI medical recommendations?- Why do nurses misclassify emergencies differently with misleading AI assistance?
- Does AI change clinician cognition or just increase reliance on predictions?
- Why do radiologists fail to benefit from AI decision support?
- Why does labeling advice as AI from a doctor change how people trust it?
- Can people tell which medical advice is accurate based only on how it reads?
- How does trusting wrong AI advice change what medical action people decide to take?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How much does wrong AI advice harm radiologist accuracy?
When mammography radiologists receive incorrect AI suggestions labeled as system output, how much does their diagnostic accuracy decline? This matters for understanding automation bias in clinical workflows.
same automation-bias pattern in radiology; this note adds a matched gain-versus-loss figure within one study
-
Does time pressure make AI advice more persuasive to experts?
When pathologists work under time constraints, does pressure to decide quickly make them more likely to trust and act on AI recommendations, even when those recommendations are wrong?
same clinical automation-bias shape; that note isolates time pressure as the severity driver, this one attributes the swing to AI changing cognition generally
-
Why don't radiologists benefit from AI predictions?
When radiologists receive AI predictions, they often fail to incorporate them properly into their decisions. This explores what belief-updating errors prevent radiologists from realizing potential AI-assisted gains.
both show average-performance metrics concealing the real effect of AI assistance on expert judgment
-
Does granting agents more autonomy undermine human oversight?
Explores whether the design of autonomous AI systems—by giving agents greater independence—actually weakens the human overseer's ability to catch problems. Matters because oversight is a key safeguard against AI failures.
shared claim that AI assistance itself, not operator failure, degrades the human's judgment or oversight capacity
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- How AI Can Degrade Human Performance in High-Stakes Settings
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- Automation Bias in AI-Assisted Medical Decision-Making under Time Pressure in Computational Pathology
- Do as AI say: susceptibility in deployment of clinical decision-aids
- Combining Human Expertise with Artificial Intelligence: Experimental Evidence from Radiology
- AI-based Clinical Decision Support for Primary Care: A Real-World Study
- Could AI Slow Science? Confronting the Production-Progress Paradox
- All AI Models are Wrong, but Some are Optimal
Original note title
Morey, Rayo and Woods find misleading AI predictions hurt nurse performance by 96 to 120 percent — more than correct predictions helped by 53 to 67 percent