Does it matter whether AI polishes people's ideas or generates them itself, and which one makes a group's thinking more alike?
Does AI refinement preserve idea diversity better than AI ideation compresses it?
This explores whether it matters *when* AI enters the creative process: does AI that polishes ideas people already had protect the variety of a group's ideas, while AI that comes up with the ideas itself makes everyone's output look alike?
This explores whether the point where AI joins the creative process changes how varied a group's ideas end up. One study answers yes. In a preregistered experiment, ideas generated by AI made the group's ideas less varied for every kind of writer, while AI that only refined ideas people had already written left that variety intact Does AI assistance homogenize or preserve creative diversity?. The most surprising result was about non-native English speakers. They contributed more varied ideas than native speakers, and AI ideation erased that advantage. AI refinement did not. When AI supplies the starting idea, people from different backgrounds lose the distinct perspectives they would otherwise bring.
The result is not settled everywhere in the collection. A critique of a different paper points out that its ideation experiment showed AI made ideas more numerous and more detailed, especially for less experienced writers, but never measured diversity Does AI assistance actually narrow the diversity of ideas?. That gap is easy to miss. More ideas and richer ideas can look like more variety when they aren't. The useful test for any claim like this is whether the study actually measured how much ideas differ from one another, or only counted them.
The reason AI-generated ideas converge shows up in research on the models themselves. A study of more than 70 models answering 26,000 open-ended prompts found that different models independently give strikingly similar answers. The authors call this the "Artificial Hivemind," and trace it to shared training data and similar alignment methods Do different AI models actually produce diverse outputs?. So switching to another model doesn't bring back the variety. A related argument holds that AI multiplies claims without multiplying viewpoints: a thousand AI-written articles may express roughly one point of view Does AI generate diverse claims or diverse perspectives?. Seen this way, refinement keeps diversity because the viewpoint still comes from the person. The AI only shapes how it is expressed.
The same divide between generating and refining appears inside model training, under different names. Reinforcement learning tends to narrow a model's behavior toward one strategy that wins the reward, while fine-tuning on varied human examples keeps its range of approaches wide Does reinforcement learning squeeze exploration diversity in search agents?. Adding step-by-step critique during training helps stop the model from settling on one approach too early Do critique models improve diversity during training itself?. Narrowing isn't automatic, though. Preference tuning makes code less varied but creative writing more varied, depending on what each domain rewards Does preference tuning always reduce diversity the same way?.
Here is the takeaway you might not have expected: variety does not come from simply putting more voices in the room. Teams of AI agents with different thinking styles produce better ideas only when the members also have real domain expertise. Without it, a diverse team does worse than a single competent agent Does cognitive diversity alone improve multi-agent ideation quality?. Diversity is only worth something when each perspective is grounded in real knowledge. That is a likely reason human-first, AI-refined work holds up better: the variety comes from people, who have actually lived the perspectives they bring.
Sources 8 notes
In a preregistered experiment, AI-generated ideas reduced collective diversity for all writers, while AI that refined existing ideas kept diversity intact. Non-native English speakers contributed more diversity than native speakers, but only AI ideation erased this advantage.
The paper's ideation experiment shows AI help increases idea count and detail, particularly for less experienced writers, but provides no diversity measure to support its conclusion about narrowed diversity.
INFINITY-CHAT analyzed 70+ models across 26K open-ended queries and found an "Artificial Hivemind" effect: models independently generate strikingly similar or identical responses due to overlapping training data and alignment procedures, undermining the diversity benefits of model ensembles.
Large language models generate numerous well-formed claims by following probabilistic patterns in training data, not by exploring competing argumentative positions. This produces volume without perspectival diversity—a thousand AI articles often represent approximately one viewpoint.
RL training compresses behavioral diversity in search agents through the same entropy collapse mechanism documented in reasoning—policies converge on narrow reward-maximizing strategies. SFT on diverse demonstrations preserves exploration breadth, suggesting diversity-preservation techniques are essential for RL search scaling.
Show all 8 sources
Step-level critique in the training loop counteracts tail narrowing and maintains solution diversity across self-training iterations. This training-time benefit—preventing premature convergence—is more fundamental than test-time accuracy gains.
RLHF reduces lexical-syntactic diversity in code generation but increases it in creative writing. The direction depends on what each domain incentivizes: code rewards convergence toward correct solutions, while creative writing rewards stylistic distinctiveness.
Multi-agent teams substantially outperform solo ideation, but only when members possess genuine senior knowledge. Diverse teams without expertise underperform even a single competent agent, because cognitive stimulation without expertise triggers process losses instead of insight.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Human diversity fuels collective creativity that large language models cannot simulate or sustain
- What and Whose Knowledge? Measuring Epistemic Diversity in Large Language Models
- NoveltyBench: Evaluating Language Models for Humanlike Diversity
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- Jointly Reinforcing Diversity and Quality in Language Model Generations
- The Homogenizing Effect of Large Language Models on Human Expression and Thought
- On Epistemic Diversity in Large Language Models
- ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs