Does AI assistance homogenize or preserve creative diversity?
Can AI tools maintain the diverse ideas that emerge from diverse human groups, or do they compress creative output toward similarity? This matters because collective diversity drives innovation.
The paper tests two threats to the idea that diverse people produce diverse ideas: everyday AI assistance may homogenize what people create, and AI-simulated diversity may replace the people altogether. Its answer to both is that "human diversity retains value that current AI can neither sustain nor simulate by default." In a preregistered creative metaphor experiment, native (L1) and non-native (L2) English writers worked without AI, with AI-generated ideas (AI ideation), or with AI refining their own ideas (AI refinement). L2 writers contributed more collective diversity than L1 writers. AI ideation compressed collective diversity for everyone and left the L2 advantage undetectable, while AI refinement preserved both. When the whole writer pool was then simulated, "every simulated pool fell below every human pool."
The framing borrows from the diversity literature the introduction cites. Country or ethnicity are treated as proxies for "deep-level diversity in values, perspectives, and cognitive repertoires," and people who represent a problem differently "search different regions of the solution space." On that account, what an AI substitute would have to reproduce is a spread of problem representations, and the simulation arm tried to supply one: personas built from participants' real backgrounds, three model families, native-language prompting, and elevated sampling temperatures. Pushing the models further produced diversity "only through degenerate text." The discussion's main design point is that "how AI enters the creative workflow decides how much of that value survives." Generating ideas for people shrinks the pool. Refining ideas people already have does not.
This extends Do different AI models actually produce diverse outputs? from model ensembles to a human benchmark: here three model families still fell short of every human pool, even with persona conditioning. It moves the individual-versus-collective split in Why do LLMs generate novel ideas from narrow ranges? onto human writers. AI ideation raised individual writers' ratings while shrinking the collective pool, which the authors describe as "pitting private incentives against the collective good." The exception was L2 writers using their native language, which benefited both. For designers of agent teams, Does cognitive diversity alone improve multi-agent ideation quality? treats diversity as something to engineer among agents. This paper reports that, for this task, engineered persona diversity did not match a real human pool. Do language models flatten the range of public arguments? gives the same reason to compare against a human distribution rather than judge a model's outputs alone.
The excerpt is silent on sample sizes, how collective diversity was measured, effect sizes, and who produced the ratings. It reports one task (creative metaphors) and one language contrast (English L1 versus L2), and it offers no mechanism for why AI ideation compresses the pool. "By default" is a real qualifier: the simulations used the settings listed above, and the excerpt does not claim no method could do better. Within those limits, the practical reading is that the position of the AI in the workflow is a design variable, and private incentives alone do not protect the collective pool, because the same intervention that lowered the pool raised individual ratings.
Inquiring lines that read this note 3
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What types of diversity prevent reasoning systems from collapsing? How should designers communicate what AI systems truly are and can do? Do writers recognize when AI writing assistance alters their expressed stance?Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do different AI models actually produce diverse outputs?
Explores whether using multiple different language models together creates genuine diversity or whether shared training and alignment cause them to converge on similar answers despite independence.
model-side convergence; this paper adds a human-pool comparison in which simulated pools fall below every human pool
-
Why do LLMs generate novel ideas from narrow ranges?
LLM research agents produce individually novel ideas but cluster them in homogeneous sets. This explores why high average novelty coexists with poor diversity coverage and what it means for automated ideation.
same individual-versus-collective split, seen here in human writers using AI ideation
-
Does cognitive diversity alone improve multi-agent ideation quality?
This explores whether diverse perspectives in group AI systems automatically produce better ideas, or if something else—like expertise—is equally critical for collaborative ideation to outperform solo agents.
engineered agent diversity contrasts with the finding that simulated pools do not match human ones
-
Do language models flatten the range of public arguments?
When LLMs write essays on the same topics as humans, do they recover the full spectrum of distinct arguments and reasons people actually make, or do they narrow the deliberative space readers encounter?
another measure of homogenization taken against a human distribution
-
Where does mode collapse in language models really come from?
Researchers investigate whether mode collapse—when models narrow to repetitive outputs—stems from training algorithms or the preference data itself. Understanding the root cause is crucial for fixing diversity loss in creative and synthetic tasks.
extends: typicality bias in preference data explains why AI ideation compresses the pool, and verbalized sampling restores diversity 1.6-2.1× without training
-
Do frontier LLMs actually explore the full space of valid answers?
When multiple correct answers exist, do advanced language models expose users to that full range, or do they collapse onto a narrow canonical subset? This matters for learning, inquiry, and decision-making.
evidence for: frontier LLMs collapse large valid answer spaces onto a few canonical answers, consistent with AI ideation shrinking the pool
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Human diversity fuels collective creativity that large language models cannot simulate or sustain
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- Has the Creativity of Large-Language Models peaked? —an analysis of inter- and intra-LLM variability —
- We Are All Creators: Generative AI, Collective Knowledge, and the Path Towards Human-AI Synergy
- NoveltyBench: Evaluating Language Models for Humanlike Diversity
- Beyond Brainstorming: What Drives High-Quality Scientific Ideas? Lessons from Multi-Agent Collaboration
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
- On Epistemic Diversity in Large Language Models
Original note title
human diversity fuels collective creativity that AI cannot simulate or sustain by default — AI ideation compresses the pool while AI refinement preserves it