Available but Unclaimed: An Empirical Study of Human-AI Synergy
People increasingly reason with large language models (LLMs), yet complementary capabilities do not guarantee outperforming both components. In a between-subjects study, participants (N=535) solved a 40-item battery of matrix reasoning, mental rotation, syllogisms, and letter-string analogies, unaided or with GPT-5.6-Luna, Claude Opus 4.8, Gemini 3.6 Flash, or Kimi K3. Each assisted trial required consultation with the model. Each model answered every item alone 100 times under matched elicitation. The assisted–unaided accuracy difference increased with item-level LLM competence. Deference varied across tasks and increased with competence within tasks. Post-advice confidence distinguished correct from incorrect answers less strongly than unaided confidence. In a reference comparison, about half the increase in LLM accuracy carried through to assisted accuracy. How much of that accuracy gain reached participants differed across the models. These findings motivate evaluating LLMs in interaction with humans and designing support for selective deference that preserves independent reasoning.
Introduction. Large language models (LLMs) are unevenly reliable within a single domain. A model may solve one reasoning problem on nearly every attempt and a problem of the same form at chance [42, 49], while its replies need not reveal the difference [83]. Whether a person can distinguish correct from incorrect advice and rely on the model accordingly matters for effective human–AI interaction (HAI). People and LLMs exhibit different strengths and weaknesses when working on a problem. One can theoretically compensate for the other’s errors [27, 72]. Complementarity denotes this potential, created where human and AI errors differ. Synergy denotes the realized case in which the team actually outperforms both components [73]. During an interaction, answers can change as the conversation proceeds. Differing errors therefore create an opportunity, but do not guarantee that the person can recognize and correct them. We call the realized share of independent-error reference headroom synergy capture. However, achieving synergy is not built into an LLM’s configuration.
Discussion / Conclusion. We asked when a person and an assistant working together outperform both components. Assisted accuracy fell below the item-wise better-component reference, while the battery-level assisted–LLM difference remained uncertain. An AI assistant is not uniformly reliable. On two items of the same kind, it can be near-certain on one and no better than guessing on the next. We measured that variation by running each assistant repeatedly on the same items. Participants’ deference varied substantially by task and also increased with item competence. Consistent with Vaccaro et al. [73], we found no clear advantage over the better component. Riedl and Weidmann’s modeled AI benefit instead compares assisted with unaided performance [63]. For AI development. These data compare solo and assisted performance on the same items. Joint evaluation has both conceptual [26] and empirical precedents [14, 63], while benchmark construct validity remains a concern [4]. Pass-through adds a measure of how assisted accuracy varies with item competence under a specified protocol.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Why can't humans reliably detect AI-generated text despite measurable linguistic signatures?- Does AI struggle with poetry for the same reason it misses jokes?
- Why do human judges fail to detect systematic linguistic differences that classifiers easily identify?
- Could AI assessment quality differ across subjects or question formats?
- Can reasoning benchmarks separate logic from believability?
- What explains the gap between perplexity performance and actual reasoning capability?
- Why do benchmark scores rise while reasoning quality declines?
- How do humans and LMs differ on multi-hop reasoning?
- Where do humans and language models actually diverge in reasoning ability?
- Do LLMs actually reason differently than humans about moral dilemmas?
- Why do LLMs fail at counterfactual reasoning despite factual knowledge?
- Why do LLMs fall for and deploy logical fallacies with equal confidence?
- Where do LLMs fail as knowledge systems compared to humans?
- How does the LLM Fallacy differ from automation bias and cognitive offloading?
- Which knowledge types do LLMs handle better than humans in reasoning tasks?
- How do LLM biases manifest differently across the three paradigms?