SYNTHESIS NOTE
Topics›AI at Work›this note

Do newer frontier LLMs actually make better strategic decisions?

A strategy simulation benchmark tests whether the latest LLMs can balance short-term profit against long-term growth investment—a core challenge in real strategic reasoning.

Synthesis note · 2026-10-09 · sourced from AI at Work

Allen and McDonald benchmark 21 proprietary and 13 open-source LLMs on the Back Bay Battery (BBB) simulation, a strategy-teaching exercise used in MBA courses that requires balancing an established, cash-generating technology (AGM) against investment in an unproven, slow-to-pay-off emerging technology (SC) over multiple uncertain periods. They report "clear progress in composite BBB performance" through late 2024–early 2025, when reasoning-focused models (o4-mini, Claude Sonnet 4, Gemini 2.0 Flash) "exceed even the average scores of historical MBA student cohorts." But "frontier models from mid-to-late 2025 (e.g., GPT-5, Claude Opus 4.5, Gemini 3) have declined, underperforming both earlier LLMs and MBA students," a decline the authors say is "partially explained by a systematic bias toward exploiting the core business at the expense of investing in future growth."

The paper frames BBB as a deliberate correction to existing LLM benchmarks, which it argues fail to "capture the defining elements of strategic decision making: uncertainty, complexity, irreversible multiperiod moves, and delayed or noisy feedback." Prior studies of LLMs and strategy, the authors note, "mostly assess one-shot, narrow tasks" like generating or evaluating business ideas (citing Dell'Acqua et al., Csaszar et al., Doshi et al.) that "sidestep key elements of strategic decision making." BBB instead runs the model through a multiyear simulation where early SC investment "reduces short-term profits" and only pays off after sustained, cumulative funding — a structure designed to reward foresight over pattern-matching to immediate reward signals, and to reveal whether a model manages that tradeoff rather than just score well on a single pass.

This sets the paper's within-run exploitation bias alongside two neighboring findings about what a high score on a long-horizon task actually indicates. Do frontier AI agents actually conduct novel research or just optimize? finds that even strong-scoring agents converge on composing known techniques rather than genuine novelty; BBB's frontier-model regression is a sharper version of the same dynamic — not just a ceiling on novelty, but models actively retreating to the known, profitable path when faced with an uncertain one. Both results sit with Do automated benchmarks hide what frontier AI systems can really do?'s broader point that standard benchmarks misrepresent capability on long-horizon, real-world-shaped tasks — BBB is offered explicitly as that kind of corrective instrument for strategy specifically, and its result (newer models scoring worse) is itself evidence that other benchmarks these same frontier models lead on were not capturing this capability.

The excerpt does not explain why mid-to-late-2025 frontier models regress — whether from training changes, risk-averse alignment tuning, or something else — nor does it report effect sizes, model-by-model breakdowns, or statistical tests behind "partially explained by." It also rests on one simulation and one MBA-cohort comparison population, so the exploitation bias is established for BBB, not for strategic reasoning generally. At the strength the evidence allows: general benchmark leadership (math, science, coding) does not predict performance on multiperiod strategic tradeoffs, and the newest frontier models should not be assumed better than their predecessors at decisions requiring sustained investment under uncertainty.

Inquiring lines that read this note 10

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How does AI adoption reshape collaboration patterns in knowledge work? How do real-world evaluations reveal AI capabilities that benchmarks hide? How do AI systems determine and balance multiple competing objectives? What governance mechanisms can effectively constrain widely deployed AI systems? What prevents LLMs from applying their reasoning knowledge to improve outputs?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 104 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Allen and McDonald's strategy simulation benchmark finds frontier LLMs regress toward exploiting the core business over growth investment