Different AI models can sound different without actually thinking differently — how would you even test for that?
How should researchers measure epistemic diversity across different language models?
This explores how you would actually test whether different AI models give a real spread of perspectives, as opposed to the same answer worded differently, and what the corpus suggests a sound measurement should include.
This explores how to check whether different language models offer a real range of viewpoints, and not just the same answer in different words. The corpus doesn't contain a methods handbook for this. It does contain several large studies that each measured diversity a different way, and together they suggest four design choices a sound measurement needs.
The first choice is a reference point, because diversity only means something when you compare it to something. One study tested 27 models across 1.7 million responses over three years. It compared them against a search-engine baseline, meaning the spread of perspectives you'd get by looking the question up on the web Are large language models becoming more epistemically diverse?. That baseline reveals something a model-to-model comparison would hide. Diversity rose substantially after 2023, but every model still fell short of search. The gains were also uneven: they depended on whether a model used retrieval, how large it was, and which language was being asked about. So a single diversity score per model hides real variation by condition. A second reference point is people. When 106 models were placed in a 'value space' built from 625 moral scenarios, the models bunched into one narrow, idealized region while human respondents spread widely Do large language models actually reflect human value diversity?. Measuring against human spread shows how far a model is from the population it might stand in for.
The second choice is to compare models with each other, not just to look at each one alone. A model can vary a lot from one answer to the next and still sound like every other model. The INFINITY-CHAT study asked 70+ models 26,000 open-ended questions and found what it calls an 'Artificial Hivemind': different models from different labs often gave strikingly similar or even identical answers. The likely cause is shared training data and similar alignment methods Do different AI models actually produce diverse outputs?. In practice, this means using several models doesn't guarantee several viewpoints. A useful measurement has to check how much models overlap with one another, not only how much each one varies internally.
The third choice is to measure the process, not just the final answers. In a study of group discussions, groups of LLMs reproduced a well-known human pattern: discussion helps average members more than the strongest ones. But they got there by conforming sooner, settling on an answer earlier, and sharing less of the information only one member had Do language model groups mimic human group reasoning patterns?. If you scored only the end result, those groups would look as diverse in their reasoning as human groups, which they weren't. The same caution applies to choosing questions. Chatbot Arena's rankings are believable partly because its question pool is varied and good at separating strong models from weak ones Can crowdsourced votes reliably rank language models?. A diversity test built on narrow or easy questions will make models look more alike than they really are.
The fourth choice, and the one you might not expect, is that the most important effect may be on the people using the models. One paper argues that because millions of people rely on the same few models, the narrowing compounds. Co-writing studies show users unknowingly adopting the model's positions and framing Do large language models narrow human expression and thought?. If that's right, a complete measurement would also check whether people's own writing and opinions become more alike after using these tools. The corpus raises this question but doesn't yet include a study that measures it.
Sources 6 notes
Testing 27 LLMs across 1.7M responses shows diversity increased substantially since 2023, yet every model remained less diverse than a search baseline. Gains varied unevenly by retrieval-augmentation, model scale, and language availability.
Analysis of 106 LLMs across 625 scenarios shows they cluster in a concentrated region of value space while human respondents scatter widely. Models are poor surrogates for diverse populations despite exhibiting coherent value systems.
INFINITY-CHAT analyzed 70+ models across 26K open-ended queries and found an "Artificial Hivemind" effect: models independently generate strikingly similar or identical responses due to overlapping training data and alignment procedures, undermining the diversity benefits of model ensembles.
LLM groups reproduce the human assembly-bonus asymmetry where discussion helps average members more than top performers, but achieve this through greater conformity, earlier convergence, and less unique information surfacing than human groups.
Chatbot Arena's 240K+ crowdsourced preference votes produce credible model rankings because the underlying questions are diverse and discriminating, and crowd judgments correlate with expert raters—validating human preference as a scalable evaluation signal.
Show all 6 sources
LLMs mirror skewed slices of human experience shaped by training data regularities, and widespread reliance on identical models amplifies convergence. Co-writing studies show users unconsciously adopt model stances and framings.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- The Homogenizing Effect of Large Language Models on Human Expression and Thought
- On Epistemic Diversity in Large Language Models
- Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
- What and Whose Knowledge? Measuring Epistemic Diversity in Large Language Models
- NoveltyBench: Evaluating Language Models for Humanlike Diversity
- Learning Pluralistic User Preferences through Reinforcement Learning Fine-tuned Summaries
- Mapping the Emerging Social Science of Large Language Models
- From Process Loss to Assembly Bonus: Human-Grounded Diagnosis of Multi-Agent LLM Collaboration