Available but Unclaimed: An Empirical Study of Human-AI Synergy

Paper · arXiv 2609.16793 · Published September 15, 2026
Reasoning Architectures

People increasingly reason with large language models (LLMs), yet complementary capabilities do not guarantee outperforming both components. In a between-subjects study, participants (N=535) solved a 40-item battery of matrix reasoning, mental rotation, syllogisms, and letter-string analogies, unaided or with GPT-5.6-Luna, Claude Opus 4.8, Gemini 3.6 Flash, or Kimi K3. Each assisted trial required consultation with the model. Each model answered every item alone 100 times under matched elicitation. The assisted–unaided accuracy difference increased with item-level LLM competence. Deference varied across tasks and increased with competence within tasks. Post-advice confidence distinguished correct from incorrect answers less strongly than unaided confidence. In a reference comparison, about half the increase in LLM accuracy carried through to assisted accuracy. How much of that accuracy gain reached participants differed across the models. These findings motivate evaluating LLMs in interaction with humans and designing support for selective deference that preserves independent reasoning.

Introduction. Large language models (LLMs) are unevenly reliable within a single domain. A model may solve one reasoning problem on nearly every attempt and a problem of the same form at chance [42, 49], while its replies need not reveal the difference [83]. Whether a person can distinguish correct from incorrect advice and rely on the model accordingly matters for effective human–AI interaction (HAI). People and LLMs exhibit different strengths and weaknesses when working on a problem. One can theoretically compensate for the other’s errors [27, 72]. Complementarity denotes this potential, created where human and AI errors differ. Synergy denotes the realized case in which the team actually outperforms both components [73]. During an interaction, answers can change as the conversation proceeds. Differing errors therefore create an opportunity, but do not guarantee that the person can recognize and correct them. We call the realized share of independent-error reference headroom synergy capture. However, achieving synergy is not built into an LLM’s configuration.

Discussion / Conclusion. We asked when a person and an assistant working together outperform both components. Assisted accuracy fell below the item-wise better-component reference, while the battery-level assisted–LLM difference remained uncertain. An AI assistant is not uniformly reliable. On two items of the same kind, it can be near-certain on one and no better than guessing on the next. We measured that variation by running each assistant repeatedly on the same items. Participants’ deference varied substantially by task and also increased with item competence. Consistent with Vaccaro et al. [73], we found no clear advantage over the better component. Riedl and Weidmann’s modeled AI benefit instead compares assisted with unaided performance [63]. For AI development. These data compare solo and assisted performance on the same items. Joint evaluation has both conceptual [26] and empirical precedents [14, 63], while benchmark construct validity remains a concern [4]. Pass-through adds a measure of how assisted accuracy varies with item competence under a specified protocol.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Why can't humans reliably detect AI-generated text despite measurable linguistic signatures? What makes specific clarifying questions more effective than generic ones? Why do benchmark improvements fail to reflect actual reasoning quality? Why do reasoning models fail at systematic problem-solving and search? Is embodied interaction necessary for language meaning and genuine agency? How do language models establish social grounding in human dialogue? How do training data properties shape reasoning capability development? How faithfully do LLMs reflect their actual reasoning in outputs and explanations? Do language models develop causal world models or rely on statistical patterns? Do accurate-looking LLM outputs hide structural failures in learning and reasoning? How do language models inherit human biases from training data? How do we evaluate AI systems when user perception misleads actual performance? Does AI fluency substitute for verifiable accuracy in human judgment? Do language models perform faithful symbolic reasoning independent of semantic grounding?