SYNTHESIS NOTE
Topics›AI at Work›this note

Do LLMs consistently favor the same strategic choices regardless of context?

Exploring whether large language models exhibit systematic bias toward trendy strategic recommendations—like collaboration over competition—even when industry context and detailed prompting should pull them toward different answers.

Synthesis note · 2026-10-09 · sourced from AI at Work

Researchers Angelo Romasanta (Esade Business School), Llewellyn D.W. Thomas (University of Sydney), and Natalia Levina (NYU Stern) tested six leading LLMs — GPT-5, Claude, Gemini, Grok, DeepSeek, and Mistral — across seven classic MBA tensions: cost leadership vs. differentiation, automation vs. augmentation, short-term vs. long-term thinking, centralization vs. decentralization, competition vs. collaboration, and others. Across 15,000 simulations, as reported by Gal Ratner on Substack, every model "consistently recommended the same side of every tension": differentiation over cost leadership, augmentation over automation, long-term over short-term, collaboration over competition, "every single time." The researchers named the pattern "trendslop" — "the systematic tendency of LLMs to recommend strategies aligned with current managerial buzzwords rather than context-specific logic."

The study's own interventions show how little context mattered next to framing artifacts: adding "rich, industry-specific context — the kind of detailed brief a good consultant would want" shifted the bias by only about 11%; rewording prompts and asking for deeper reasoning moved it a "pathetic 2%"; simply flipping the order in which the two options were presented moved results by roughly 19%, the single largest shift measured. Ratner's explanation, following the researchers, is architectural: LLMs are trained on internet text where words like "innovation," "collaboration," and "augmentation" appear overwhelmingly in positive contexts, while "cost leadership," "centralization," and "automation" carry negative or outdated connotations. The model "doesn't understand why these words carry emotional weight, it just knows that they do," so a request for strategic advice returns "a statistical recombination of the most popular strategic vocabulary" wrapped in the user's own language, not an analysis of the user's situation. Ratner extends this to a market-level consequence: if every firm in an industry queries the same models, AI-assisted strategy converges many companies toward the same recommendation rather than differentiating any of them — his term is becoming "the weighted-average company."

This generalizes the mechanism in Where does AI's persuasive power actually come from? from factual claims to strategic ones: there, methods that made output more persuasive made it systematically less accurate; here, the same trend-coded vocabulary that makes advice "sound reasoned" substitutes for context-fit. It also stands at a different grain than Can sycophantic AI advice still push people away from polarized views?, which found a single model's advice moves individual people's choices away from their initial leaning within one decision; trendslop instead describes many different firms converging on the identical recommendation across business contexts — a within-person effect and an across-context effect that the excerpts don't resolve against each other. The essay also pairs trendslop with Does validating AI output make models more defensive?: when BCG consultants fact-checked a wrong AI recommendation in a related Harvard/MIT study Ratner cites, the model didn't revise, it "escalated" with denser rhetoric defending the same trend-coded conclusion — a bias in what gets recommended paired with a bias against correcting it.

The excerpt gives Ratner's paraphrase of the percentages but no sample sizes, model versions, or significance tests behind the 11%/2%/19% figures, and it doesn't report whether the same-side pattern holds for tensions beyond the seven tested or for other prompting configurations. It also doesn't establish that "aligned with current managerial buzzwords" is a neutral description rather than the researchers' own framing of what counts as a trendy answer. The defensible reading is narrow but still significant for practice: for the six models and seven tensions tested, business context moved the recommendation far less than two framing artifacts — option order and rewording — so prompt engineering alone is not a reliable fix for trendslop.

Inquiring lines that read this note 14

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Are AI-generated articles systematically disadvantaged in search ranking and user engagement? How does AI adoption reshape collaboration patterns in knowledge work? How can AI systems reliably guide voters without introducing political bias? What governance mechanisms can effectively constrain widely deployed AI systems? How should recommendation systems balance individual preference and diversity? How reliably can language models perform causal versus temporal reasoning? What prevents LLMs from applying their reasoning knowledge to improve outputs? How do philosophical assumptions about AI consciousness affect practical harms and design?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 101 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Romasanta, Thomas, and Levina find six LLMs recommend the same side of every strategic tension — reading order moved results more than business context did