Human diversity fuels collective creativity that large language models cannot simulate or sustain

Paper · arXiv 2607.26899 · Published July 29, 2026
LLM Failure Modes

Diverse human groups produce diverse ideas, the raw material of innovation. Generative AI challenges this engine twice over: everyday AI assistance may homogenize what diverse people create, and AI-simulated diversity may replace the people altogether. We tested both challenges in a preregistered creative metaphor experiment with native (L1) and non-native (L2) English writers, who wrote without AI, with AI-generated ideas (AI ideation), or with AI refining their own ideas (AI refinement). L2 writers contributed more collective diversity than L1 writers, with native-language ideation showing the most diverse pools. AI ideation compressed collective diversity for everyone and left the L2 advantage undetectable, whereas AI refinement preserved both. We then simulated the entire writer pool using personas built from participants’ real backgrounds, three model families, native-language prompting, and elevated sampling temperatures. Every simulated pool fell below every human pool, and pushing models further induced diversity only through degenerate text. However, at the individual level, AI ideation raised writers’ ratings, pitting private incentives against the collective good, except when L2 writers used their native language, which benefited both.

Introduction. Diverse human composition has long been considered a source of creativity and innovation. In scientific research, teams that are diverse in gender, ethnicity, and location produce more novel and higher-impact work (AlShebli et al., 2018; Freeman & Huang, 2015; Yang et al., 2022). Meta-analytic evidence from organizational research echoes this pattern, linking cultural diversity to team creativity and innovation (Wang et al., 2019). Notably, these benefits do not require face-to-face collaboration. Groups of diverse problem solvers outperform groups of high- ability problem solvers even when members contribute independently (Hong & Page, 2004). Why does who creates shape what gets created? The benefits of diversity are not fully explained by surface-level attributes such as country or ethnicity. Instead, these attributes are proxies for deep-level diversity in values, perspectives, and cognitive repertoires (van Knippenberg & Schippers, 2007). People who represent a problem differently search different regions of the solution space (Hong & Page, 2004).

Discussion / Conclusion. Human diversity has long been treated as an engine of collective creativity. We asked whether that engine still matters in the age of AI, testing the two challenges it now faces: (1) everyday LLM assistance may erode the collective diversity that diverse humans contribute, and (2) LLM- simulated diversity may replace the humans altogether. Our answer to both is that human diversity retains value that current AI can neither sustain nor simulate by default. Importantly, how AI enters the creative workflow decides how much of that value survives. In our experiment, linguistically diverse (L2) writers contributed more collective diversity than native (L1) English writers, an advantage that AI refinement preserved but AI ideation left undetectable; pools were numerically most diverse when L2 writers ideated in their native language. AI ideation also shrank everyone’s pool, whereas AI refinement preserved it.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can AI systems develop genuine social understanding without embodiment? Can AI-generated outputs constitute genuine knowledge or valid claims? How do evaluation biases undermine LLM quality assessment systems? Why does reinforcement learning suppress output diversity compared to supervised fine-tuning? Why should disagreement be treated as signal in collaborative reasoning? Can ensemble evaluation methods reduce bias more than single judges? How do multi-agent systems achieve genuine cooperation and reasoning? When should tasks involve human-AI partnership versus full automation? Can debate mechanisms prevent silent agreement on wrong answers in multi-agent reasoning? Can prompting inject entirely new knowledge into language models? How can AI systems learn from failures without cascading errors? When does optimizing for quality undermine the value of diversity? When do multi-agent approaches outperform single model extended thinking?