SYNTHESIS NOTE
Topics›Reasoning Architectures›this note

Why does assisted accuracy capture only half the LLM gain?

When an AI system improves on a task, how much of that improvement actually reaches people using it with the AI? Understanding this gap matters because it shows whether complementary strengths automatically translate to better team performance.

Synthesis note · 2026-09-25 · sourced from Reasoning Architectures

"Available but Unclaimed" separates two things that are easy to run together. Complementarity is the potential created where human and AI errors differ; synergy is the realized case in which the team outperforms both components. In a between-subjects study (N=535) on a 40-item battery of matrix reasoning, mental rotation, syllogisms, and letter-string analogies, the paper reports that the assisted–unaided accuracy difference increased with item-level LLM competence, and that in a reference comparison "about half the increase in LLM accuracy carried through to assisted accuracy." Assisted accuracy fell below the item-wise better-component reference, and the battery-level assisted–LLM difference "remained uncertain." Its summary line is that "achieving synergy is not built into an LLM's configuration."

The measurement design is what makes this readable. Each of four models answered every item alone 100 times under matched elicitation, so each item carries its own competence estimate, and each assisted trial required consultation with the model. The paper's premise is that an assistant is "not uniformly reliable": on two items of the same kind it can be near-certain on one and no better than guessing on the next, and its replies need not reveal which is which. The paper names the realized share of independent-error reference headroom "synergy capture," and adds pass-through as a measure of how assisted accuracy varies with item competence under a specified protocol.

The behavioral results sit alongside this. Deference varied across tasks and increased with competence within tasks, so participants leaned on the model more where it was better. Post-advice confidence distinguished correct from incorrect answers less strongly than unaided confidence did. The paper's own framing is that differing errors create an opportunity but do not guarantee the person can recognize and correct them.

Against Do users worldwide trust confident AI outputs even when wrong?, this is a contrast in scope. That note describes users following confident outputs even when wrong; here deference at least rises with item competence, so the failure looks less like blanket overreliance and more like partial, imperfect selectivity. The unevenness the paper starts from also fits Why do language models fail confidently in specialized domains?, where model confidence is a poor signal of model accuracy.

The excerpt does not establish why pass-through falls short. It reports the weakened post-advice confidence signal and the rising deference next to the partial pass-through but does not say that one explains the other. It also gives no per-model figures beyond noting that the share of accuracy gain reaching participants "differed across the models," and no numbers for deference or confidence. It supports two implications at the paper's own strength: an LLM's solo benchmark accuracy overstates what a person working with it will achieve, so models should be evaluated "in interaction with humans," and support should be designed for "selective deference that preserves independent reasoning."

Inquiring lines that read this note 7

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does AI assistance promote real skill development or substitute for independent learning? How do capability benchmark scores systematically misrepresent true model abilities? How can infrastructure records verify actual agent behavior? What training dynamics and scale trigger emergence of reasoning capabilities? What fundamental constraints limit how effectively agents can improve themselves?

Related concepts in this collection 2

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 117 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

assisted accuracy carries through only about half of the gain in item-level LLM accuracy, so complementarity does not guarantee synergy