INQUIRING LINE

Are AI test scores missing the real failures, where people take, ignore, or over-trust what the AI tells them?

Do standard benchmarks miss how humans actually fail to use AI advice?

This explores whether the usual way of scoring AI (accuracy on a test set) misses where human-AI work actually breaks down: in how people take, ignore, or over-trust what the AI tells them.


This explores whether the usual way of scoring AI (accuracy on a test set) misses where human-AI work actually breaks down: in how people take, ignore, or over-trust what the AI tells them. The corpus says yes, though it gets there from the side. Few notes here study people using AI advice directly. Instead, several of them show that a benchmark scores the model alone, while the trouble shows up when a person acts on what the model said.

The clearest link is about which mistakes do the damage. Accuracy averages over everything, so a model can look strong while its fluent, confident errors pile up in the rare cases where harm happens, like an unusual medical triage case, an odd legal clause, or a financial plan with an unstated constraint Why do confident wrong answers hide in standard accuracy metrics?. Those are exactly the cases where a person is least able to catch the error, because the answer sounds just as sure as the correct ones. The Rose-Frame work explains why people miss it. Treating the output as if it were the thing itself, mistaking quick intuition for careful reasoning, and having existing beliefs confirmed are three traps that make each other worse. Fluency invites trust that accuracy scores never measure Why do people trust AI outputs they shouldn't?.

The more surprising point is that how advice is delivered may matter as much as whether it's right. When a model hands over a finished answer, people anchor on it. Learning to Guide reverses this: the AI points out which parts of the input deserve attention and leaves the decision to the person. In those studies, that removed the anchoring problem Can AI guidance reduce anchoring bias better than AI decisions?. A benchmark that scores only the answer can't tell these two designs apart, even though they lead to very different human behavior. Magentic-UI reaches a similar conclusion from the agent side. Nobody has a ground truth for when an AI should hand off to a human, so the system spreads that judgment across several touchpoints: planning together, guarding risky actions, and verifying steps human-ai-collaborative-systems-require-six-interaction-mechanisms-because-the. Part of the failure is the model not knowing what it doesn't know about the person. Simply listing those unknowns in the prompt sharply reduced harmful advice and sycophancy Do language models know what they don't know about users?.

Zoom out and the same pattern appears in general critiques of benchmarks. Tests favor tasks that are precisely specified and easy to grade automatically, so they overstate some abilities and understate others Do automated benchmarks hide what frontier AI systems can really do?. Agents win contest-style tasks but stall on real occupational workflows Why do agent benchmarks not predict real economic value?. In each case the field optimizes what it measures, and it has measured the model on its own rather than the person and the AI working together.

A direct caveat: this corpus doesn't have controlled studies that put people in front of AI advice and measure how often they wrongly accept or reject it. The argument above is built from model-side evaluation critiques and interaction-design work. What it does suggest is a useful reframe: "how accurate is the AI?" and "does this person make better decisions with it?" are different questions, and only the first one is on the leaderboard.


Sources 7 notes

Why do confident wrong answers hide in standard accuracy metrics?

Medical triage, legal interpretation, and financial planning show a consistent pattern: surface heuristics conflict with unstated constraints, producing fluent confident errors that concentrate in rare cases where harm occurs. Aggregate accuracy masks these failures because overall performance looks strong.

Why do people trust AI outputs they shouldn't?

Rose-Frame identifies map-territory confusion, intuition-reason conflation, and confirmation-bias reinforcement as traps that multiply their distorting effects when they co-occur. Evidence from cross-linguistic overreliance and architectural transformer biases confirms the compounding mechanism operates universally.

Can AI guidance reduce anchoring bias better than AI decisions?

Learning to Guide eliminates anchoring bias and unassisted hard cases by having machines supply interpretive guidance rather than autonomous decisions, keeping responsibility with humans while improving their judgment through enhanced perception.

When should human-agent systems ask for human help?

Magentic-UI identifies co-planning, co-tasking, action guards, verification, memory, and multitasking as mechanisms that work around the lack of ground truth for optimal deferral timing. Rather than solving the timing problem directly, these mechanisms distribute decision-making across multiple touchpoints.

Do language models know what they don't know about users?

Research shows assistants suffer from sycophancy and hallucination because they have no representation of what remains unknown about users. Adding a schema of labeled unknowns to prompts reduced harmful advice and sycophancy by 50–75% and cut hallucination rates by roughly half.

Show all 7 sources
Do automated benchmarks hide what frontier AI systems can really do?

Automated benchmarks both overstate and understate capability by privileging precisely-specified, auto-gradable tasks. Open-world evaluations of long-horizon messy tasks through qualitative log analysis—with cost explicitly reported—correct these distortions and catch emerging capabilities earlier.

Why do agent benchmarks not predict real economic value?

ALE's analysis of 960 real occupational workflows shows agents excel at abstract contests but fail long-horizon professional tasks. The gap is not model capability but benchmark design—the field optimizes what it measures, and it has measured contests rather than work.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.