INQUIRING LINE

When a person and AI judge together, the pair often lands below the AI alone, and people keep only about half of its gains.

How do assisted accuracy rates compare when humans work together with AI models?

This explores what happens to accuracy when a person and an AI model make judgments together: does the pair beat either one alone, and why does it so often fall short?


This explores what happens to accuracy when a person and an AI model work on the same judgment, and whether the pair ends up better than either one alone. The most direct evidence in the collection is sobering. In a study of 535 participants, when the model got better on a given item, the people it was assisting kept only about half of that improvement. Their combined accuracy often landed below what the stronger partner, the AI, could have scored by itself Why does assisted accuracy capture only half the LLM gain?. The key distinction is that human and AI having different strengths (complementarity) only creates the possibility of a better team result (synergy). It doesn't guarantee one. Much of the AI's advantage gets lost when it passes through a human decision.

Why does it get lost? One answer is in how the AI's help is delivered. When the AI hands over a finished answer, people tend to anchor on it: they defer when they shouldn't, or they push back unevenly, and the hard cases still fall to the human without much support. One alternative flips the setup. Instead of the machine deciding and the human accepting or overriding, the machine points out which parts of the input deserve attention, and the human makes the call. This approach, called Learning to Guide, removes the anchoring problem and keeps responsibility with the person while still improving their judgment Can AI guidance reduce anchoring bias better than AI decisions?. The lesson is that the half-gain result may come less from people ignoring good advice and more from the way the advice is framed.

There is also a cost that per-answer accuracy scores don't show. AI suggestions can hurt reasoning even when they are correct, because they break a person's concentration and force them to rebuild their train of thought before continuing Does AI assistance always help reasoning or does it carry hidden costs?. So a correct suggestion can lower the team's overall performance. If you only measure whether each suggestion was right, you miss what it did to the rest of the task.

The AI side of the partnership has weak spots too. People can only calibrate how much to trust the model if its confidence means something. One line of work shows that a model's confidence becomes more reliable when it is grounded in its own record of past successes and failures on similar cases, rather than in the current answer alone Can past performance predict when a model will be right?. Assistants also have no built-in sense of what they don't know about the person they're helping. Adding an explicit list of these unknowns to the prompt cut sycophantic and harmful advice by 50–75% Do language models know what they don't know about users?. Both problems make it harder for a person to tell when to lean on the AI and when to override it.

The collection has only one study that directly measures human-AI team accuracy, so treat the "about half" figure as a striking single data point, not a settled law. Still, the takeaway holds up: the shortfall in a human-AI team usually traces back to the handoff. That covers how advice is framed, when it interrupts, and whether the AI's confidence can be trusted, more than either partner's raw ability.


Sources 5 notes

Why does assisted accuracy capture only half the LLM gain?

A 535-participant study found that when LLM accuracy improved on individual items, assisted participants captured roughly half that gain—falling below what the better-performing component could have provided alone. This shows complementarity creates potential but does not guarantee synergy.

Can AI guidance reduce anchoring bias better than AI decisions?

Learning to Guide eliminates anchoring bias and unassisted hard cases by having machines supply interpretive guidance rather than autonomous decisions, keeping responsibility with humans while improving their judgment through enhanced perception.

Does AI assistance always help reasoning or does it carry hidden costs?

Well-intentioned AI suggestions can damage reasoning performance by severing cognitive immersion, forcing users to rebuild focus before continuing. Evaluation must measure flow preservation across entire tasks, not just local suggestion accuracy.

Can past performance predict when a model will be right?

XConf matches ten-sample self-consistency at a tenth of the cost by retrieving the model's past episodes with similar confidence levels and reading their historical success rates. Ablations show the signal depends entirely on stored outcomes, not on the retrieval prompt itself.

Do language models know what they don't know about users?

Research shows assistants suffer from sycophancy and hallucination because they have no representation of what remains unknown about users. Adding a schema of labeled unknowns to prompts reduced harmful advice and sycophancy by 50–75% and cut hallucination rates by roughly half.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.