Does using LLMs actually improve strategic decision making?
An experiment tested whether LLM assistance changes how people think through strategic choices and whether those changes lead to better predictions. Understanding this matters because organizations increasingly rely on AI to augment human decision-making.
Kanis, Mann, and Stumpf-Wollersheim ran a 2×2 between-participants experiment (time constraints: yes/no × LLM use: yes/no; N=348) on a startup-evaluation task adapted from prior strategic-foresight research: participants watched two competing Kickstarter pitch videos, listed weighted pros and cons for each, then picked the startup they expected to succeed and rated each one's likelihood of success. Participants in the LLM condition used a GPT-4o chatbot, prompted to supply weighted pros and cons, during the evaluation. The headline result, in the authors' words: "both time constraints and LLM use significantly alter the characteristics of mental representations," yet "neither time constraints nor LLM use are found to significantly change strategic foresight." They call this "a cautionary case for the effectiveness of LLM use in strategic decision-making."
The mechanism runs through three characteristics of mental representations defined in prior work (Csaszar and Laureiro-Martínez 2018): breadth (categorical diversity of cues used), depth (within-category detail), and consensus (similarity to the crowd's representation). Time constraints "selectively reduce breadth without significantly affecting depth" and push representations toward higher consensus, which the authors read as convergence on "similar, local, and salient cues" under pressure. LLM use moves representations the other way — notably, it "increased the share of non-consumer items in participants' mental representations under time constraints," surfacing distant cues time pressure would otherwise hide. But "additional analyses indicate … that LLM use increases information overload and reduces psychological ownership," echoing prior findings the paper cites (Draxler et al. 2024 on ownership; Kosmyna et al. 2025 on overload). Representations changed in both directions; predictive accuracy did not move with them. The authors' conclusion is that "whether mental representations improve strategic foresight could depend on how representations are generated" — richer or more consensual inputs are not automatically better-decision inputs.
That decoupling of representation from decision quality is the same gap formalized, in a different register, by Why do accurate predictions lead to poor decisions?: a model (or in this case, a human decision-maker's mental model) can be optimized to fit available cues without that fit translating into better downstream decisions, because the objective that shapes the representation is agnostic to the objective that judges the decision. Here that gap is demonstrated behaviorally rather than proven formally. The overload and ownership-loss findings also instantiate, in a workplace-strategy setting, what Why do people trust AI outputs they shouldn't? frames as a governance problem: Rose-Frame argues System 2 oversight must govern LLM-assisted System 1 fluency, and this experiment supplies a case where more cues plus less ownership is exactly the failure mode such governance would need to catch, since the extra breadth did not convert into measurably better foresight.
The excerpt does not report how information overload and psychological ownership were measured (self-report scale versus behavioral trace), so it isn't possible to say here whether "increases overload" is a reported experience or an observed effect on performance — a distinction the paper itself may draw in its measures section, which is not in this excerpt. The task is a single Kickstarter-pitch evaluation with a single model (GPT-4o, November 2026 snapshot) and an unspecified online-experiment population, not described here as practicing managers; the authors themselves flag generalizability as a limitation and call for testing the same design in "mergers and acquisitions, new market entries, or sustainability transitions." The finding to carry forward is narrow but real: in this task, giving people an LLM changed what they attended to without making their predictions better, and the paper gives no evidence this would reverse with a different model, task, or more experienced population — only that it did not happen here.
Inquiring lines that read this note 7
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do AI systems determine and balance multiple competing objectives? How does AI adoption reshape collaboration patterns in knowledge work? How can we detect and account for LLM involvement in academic writing? What prevents LLMs from applying their reasoning knowledge to improve outputs?Related concepts in this collection 2
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Why do accurate predictions lead to poor decisions?
Predictive models are built to fit data, not to optimize decision outcomes. This note explores when and why accurate forecasts fail to produce good choices.
same prediction/decision-quality gap, here demonstrated behaviorally in human strategic cognition rather than proven formally
-
Why do people trust AI outputs they shouldn't?
When do human cognitive shortcuts fail in AI interaction? Three compounding traps—treating statistical patterns as facts, mistaking fluency for understanding, and avoiding disagreement—may explain systematic overreliance across languages and contexts.
the overload and ownership-loss findings are a concrete instance of the governance failure Rose-Frame diagnoses, in a strategy-task setting
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- How Well Can AI Do Strategy? Empirical Benchmarking Using Strategy Simulations
- AI-Augmented Strategic Decision-Making Under Time Constraints: An Experimental Study on Mental Representations and Strategic Foresight
- Determinants of LLM-assisted Decision-Making
- Your AI Strategy Advisor Is Giving Everyone the Same Advice
- AI Meets the Classroom: When Does ChatGPT Harm Learning?
- Can Machines Think Like Humans? A Behavioral Evaluation of LLM-Agents in Dictator Games
- LLM Strategic Reasoning: Agentic Study through Behavioral Game Theory
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
Original note title
LLM use changes mental representations in strategic decision-making without improving strategic foresight — and increases overload while reducing ownership