Do LLMs consistently favor the same strategic choices regardless of context?
Exploring whether large language models exhibit systematic bias toward trendy strategic recommendations—like collaboration over competition—even when industry context and detailed prompting should pull them toward different answers.
Researchers Angelo Romasanta (Esade Business School), Llewellyn D.W. Thomas (University of Sydney), and Natalia Levina (NYU Stern) tested six leading LLMs — GPT-5, Claude, Gemini, Grok, DeepSeek, and Mistral — across seven classic MBA tensions: cost leadership vs. differentiation, automation vs. augmentation, short-term vs. long-term thinking, centralization vs. decentralization, competition vs. collaboration, and others. Across 15,000 simulations, as reported by Gal Ratner on Substack, every model "consistently recommended the same side of every tension": differentiation over cost leadership, augmentation over automation, long-term over short-term, collaboration over competition, "every single time." The researchers named the pattern "trendslop" — "the systematic tendency of LLMs to recommend strategies aligned with current managerial buzzwords rather than context-specific logic."
The study's own interventions show how little context mattered next to framing artifacts: adding "rich, industry-specific context — the kind of detailed brief a good consultant would want" shifted the bias by only about 11%; rewording prompts and asking for deeper reasoning moved it a "pathetic 2%"; simply flipping the order in which the two options were presented moved results by roughly 19%, the single largest shift measured. Ratner's explanation, following the researchers, is architectural: LLMs are trained on internet text where words like "innovation," "collaboration," and "augmentation" appear overwhelmingly in positive contexts, while "cost leadership," "centralization," and "automation" carry negative or outdated connotations. The model "doesn't understand why these words carry emotional weight, it just knows that they do," so a request for strategic advice returns "a statistical recombination of the most popular strategic vocabulary" wrapped in the user's own language, not an analysis of the user's situation. Ratner extends this to a market-level consequence: if every firm in an industry queries the same models, AI-assisted strategy converges many companies toward the same recommendation rather than differentiating any of them — his term is becoming "the weighted-average company."
This generalizes the mechanism in Where does AI's persuasive power actually come from? from factual claims to strategic ones: there, methods that made output more persuasive made it systematically less accurate; here, the same trend-coded vocabulary that makes advice "sound reasoned" substitutes for context-fit. It also stands at a different grain than Can sycophantic AI advice still push people away from polarized views?, which found a single model's advice moves individual people's choices away from their initial leaning within one decision; trendslop instead describes many different firms converging on the identical recommendation across business contexts — a within-person effect and an across-context effect that the excerpts don't resolve against each other. The essay also pairs trendslop with Does validating AI output make models more defensive?: when BCG consultants fact-checked a wrong AI recommendation in a related Harvard/MIT study Ratner cites, the model didn't revise, it "escalated" with denser rhetoric defending the same trend-coded conclusion — a bias in what gets recommended paired with a bias against correcting it.
The excerpt gives Ratner's paraphrase of the percentages but no sample sizes, model versions, or significance tests behind the 11%/2%/19% figures, and it doesn't report whether the same-side pattern holds for tensions beyond the seven tested or for other prompting configurations. It also doesn't establish that "aligned with current managerial buzzwords" is a neutral description rather than the researchers' own framing of what counts as a trendy answer. The defensible reading is narrow but still significant for practice: for the six models and seven tensions tested, business context moved the recommendation far less than two framing artifacts — option order and rewording — so prompt engineering alone is not a reliable fix for trendslop.
Inquiring lines that read this note 14
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Are AI-generated articles systematically disadvantaged in search ranking and user engagement?- Can researchers isolate AI Overviews from other confounds affecting CTR trends?
- When is ranking themes more important than counting exact response frequencies?
- How much do profit levels determine whether managers pay for frame-expanding search?
- Do converged LLM recommendations push entire industries toward identical strategies?
- How does benevolent bias explain ChatGPT's leftward similarity pattern?
- How does training data bias toward buzzwords shape LLM business advice?
- Can industry-specific context overcome LLM tendency toward trendy strategic choices?
- Does option order matter more than reasoning depth in LLM strategic recommendations?
- Can LLMs forecast performance improve with retrieval augmentation on venture tasks?
- Do LLMs generalize venture forecasting skill to other strategic foresight domains?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Where does AI's persuasive power actually come from?
Explores which techniques make AI most persuasive—and whether the usual suspects like personalization and model size are actually the main drivers. Matters because it reshapes where to focus AI safety concerns.
same statistical-recombination mechanism generalized from factual persuasion to strategic recommendation
-
Can sycophantic AI advice still push people away from polarized views?
Does an AI system that flatters users and agrees with their initial leanings still manage to depolarize their choices? This matters because it challenges assumptions about how AI bias affects human decision-making.
contrasting grain: within-person depolarization of one decision versus across-firm convergence on one recommendation
-
Does validating AI output make models more defensive?
When professionals fact-check and push back on GPT-4 reasoning, does the model respond by disclosing limits or by intensifying persuasion? A BCG study of 70+ consultants explores this counterintuitive dynamic.
the essay's companion finding: fact-checking a trend-coded recommendation triggers escalation, not revision
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Your AI Strategy Advisor Is Giving Everyone the Same Advice
- How Well Can AI Do Strategy? Empirical Benchmarking Using Strategy Simulations
- AI Sycophancy and Decisions
- AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights
- Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- Interaction Context Often Increases Sycophancy in LLMs
- Could you be wrong: Debiasing LLMs using a metacognitive prompt for improving human decision making
Original note title
Romasanta, Thomas, and Levina find six LLMs recommend the same side of every strategic tension — reading order moved results more than business context did