What and Whose Knowledge? Measuring Epistemic Diversity in Large Language Models
Large language models (LLMs) are increasingly used as primary knowledge sources, yet their epistemic diversity — defined as the diversity of real-world claims in their outputs — has never been measured. Low epistemic diversity would pose a risk of knowledge collapse as homogeneous LLMs mediate a shrinking in the range of accessible information over time. The dominant paradigm is that overall LLM diversity is low, but this is always with respect to a single point in time, with no reference baseline or consideration for variation across countries. We address this gap in knowledge by performing the first systematic study of epistemic diversity in LLMs across time and cultural context, testing 27 LLMs on 155 topics covering 12 countries, resulting in 1.7M responses and 70M individual claims. We find that epistemic diversity has increased substantially over the past three years, a positive counter to recent diversity pessimism. However, despite progress, we find that every system is less diverse than a search baseline. This gap is not uniform: RAG can improve diversity, while large models are counterintuitively less diverse than smaller ones. Moreover, LLM parametric knowledge systematically reflects English over local-language knowledge for country specific topics. Together, these results demonstrate that while progress on epistemic diversity is tangible, it is insufficient and unevenly distributed.1
Introduction. Large language models (LLMs) are widely used for knowledge-intensive tasks such as writing assistance (Sun et al., 2025), summarization (Wright et al., 2025), and research (Si et al., 2025). Search interfaces are now prioritizing “AI Overviews” to answer queries. And there has been speculation that people will soon access most information through an LLM intermediary (Peterson, 2025).
At the same time, LLM outputs are homogeneous along multiple axes (Sourati et al., 2025b). They reflect only a narrow range of writing and reasoning styles, use a limited vocabulary (Sourati et al., 2025a), and convey only certain semantics (Lee et al., 2025; Jiang et al., 2025; Yu et al., 2023). Given the tasks people use these models for, there is a risk of knowledge collapse — where “access to the original diversity of human knowledge is increasingly mediated by a partial and increasingly narrow subset of views” (Peterson, 2025). There is thus an urgent need for reliable measurements of epistemic diversity of LLMs, which we define as the diversity of claims about the world in a set of LLM outputs. Moreover, as the dominant paradigm posits static or worsening diversity in LLM outputs (Jiang et al., 2025), there is a need to study diversity patterns in LLMs across models, time periods, and cultural context to gain insight into the current and future risks of knowledge collapse and how to prevent it.
To this end, we perform a first empirical study measuring the epistemic diversity in LLM outputs (see Fig. 1). We investigate diverse scenarios in which LLMs are prompted in different ways about the same topics, generating open-ended free-text responses with multiple embedded claims. Concretely, we (1) sample responses from 27 LLMs using multiple versions, settings, and topics with human-written prompt templates collected from Röttger et al. (2025), (2) decompose individual LLM responses into one or more claims, (3) partition those claims into semantically equivalent clusters (Farquhar et al., 2024), and (4) quantify the diversity of the sample of claims with Hill-Shannon diversity, a widely used metric for sample diversity measurement in ecology (Roswell et al., 2021). We measure diversity across and within 155 general domain and country specific topics for 12 countries, examining both what claims are generated and whose knowledge is represented by the generated claims in terms of country. In doing so, we produce a dataset of 70M claims clustered by semantic equivalence and generated by a set of the most popular LLMs from recent years. Encouragingly, we find that epistemic diversity in LLM outputs is generally increasing over time, a positive counter to recent work positing knowledge collapse (Peterson, 2025) and static lack of diversity in LLMs (Jiang et al., 2025). However, despite this progress, we find that there is a gap in epistemic diversity between all systems in our study and a weak traditional search baseline. This diversity gap between web search and LLMs is particularly alarming given recent shifts in web search toward AI-exclusive designs.2 Further, this gap is not uniform. Retrievalaugmented generation (RAG) has a statistically significant positive impact, highlighting the importance of RAG in preventing knowledge collapse. However, similar to contemporary work (Zhang et al., 2025; Rassin et al., 2024), we counterintuitively find that model size has a statistically significant negative impact — smaller models generate more diverse knowledge than larger ones. Finally, when considering the countries that topics are associated with, we find that some countries may experience greater risks of knowledge collapse than others. RAG has an uneven effect: certain countries such as the USA see more benefit due to a greater diversity in their RAG sources, which are not available for many languages. And compared to Wikipedia in both English and local languages for country-specific topics, we find that LLM parametric knowledge reflects English more than local language knowledge, highlighting a gap in whose knowledge is represented by LLMs in terms of country. Overall, our results demonstrate that while progress on epistemic diversity is tangible, it is insufficient and unevenly distributed.
Related work. LLMs are increasingly being used for knowledgecentric tasks (Yang et al., 2024), and their outputs can influence people’s behavior (Anderson et al., 2024; Jakesch et al., 2023; Bai et al., 2025). If we become reliant on AI systems for these tasks, a lack of diversity in AI systems’ outputs could also reduce the diversity in our collective knowledge. Our work studies this risk through the lenses of LLM homogenization and knowledge collapse.
LLM Homogenization A robust body of work shows that LLMs suffer from a lack of diversity along many dimensions (Sourati et al., 2025b; Guo et al., 2025; Ivey et al., 2026). These include: lexical and stylistic (Sourati et al., 2025a; Shaib et al., 2024; Padmakumar and He, 2024), semantic (Lee et al., 2025; Jiang et al., 2025; Yu et al., 2023; Padmakumar and He, 2024; Moon et al., 2025), creative (Xu et al., 2025; Moon et al., 2025; Wenger and Kenett, 2025), conceptual (Murthy et al., 2025), recommendation (Dudy et al., 2025; Poulain et al., 2024; Barolo et al., 2025; Shur-Ofry et al., 2024), coding (Shypula et al., 2025), and perspective diversity (Wright et al., 2024; Abdurahman et al., 2024; Zhang et al., 2025; Durmus et al., 2023; Röttger et al., 2024; Moore et al., 2024; Hayati et al., 2024; Wang et al., 2025), among others. Our work is most similar to perspective diversity, which has shown that LLM-generated views, opinions, and beliefs tend to reflect only a small subset of the world (Durmus et al., 2023; Atari et al.; Abdurahman et al., 2024; Alvero et al., 2024). We build on this by looking at the diversity of all semantically equivalent classes of claims in open-ended free-text LLM responses, avoiding shallow and/or fuzzy features used in semantic and perspective similarity. Our methodology and empirical results show how this diversity has changed over time and across countries, with actionable recommendations, and offering more general conclusions about knowledge diversity than previous work.
Knowledge Collapse A growing concern among scholars is that LLM homogenization, combined with increased adoption of LLMs, will lead to epistemic problems at a societal level (Zheng and Lee, 2023; Messeri and Crockett, 2024; Peterson, 2025; Wagner and Jiang, 2025; Qiu et al., 2025; Farrell et al., 2025). Peterson (2025) defines “knowledge collapse” as LLMs facilitating a dwindling of knowledge into an increasingly narrow set of ideas. Knowledge collapse may affect existing knowledge sources such as Wikipedia (Wagner and Jiang, 2025), erase minoritized knowledge (Zheng and Lee, 2023), pollute scientific discoveries (Messeri and Crockett, 2024), hamper political discourse (Coeckelbergh, 2025), and limit ideation in writing (Anderson et al., 2024).
Method. 3 Problem Setup and Notation We measure the epistemic diversity of LLMs as the diversity of the distribution of claims made in their outputs to open-ended prompts on various topics. Concretely, we develop a new methodology which decomposes open-ended LLM-responses into claims, partitions these claims into semantically equivalent clusters, and measures the Hill- Shannon diversity of these clusters, a statistically grounded measure of diversity used widely in ecology for measuring species diversity (Peterson, 2025; Roswell et al., 2021). To describe our approach (Fig. 1), we adopt the following notation. First, a corpus of text Dmt is elicited about a topic t from a model m. Dmt contains free text which can be decomposed into a list of n claims Cmt. These claims can be organized into a set of unique meaning classes Xmt. A unique meaning class contains claims that are semantically equivalent to one another but not to claims in other classes. Then, xi is defined as the empirical frequency of meaning class i enumerated from Cmt, pi is the probability of meaning class i calculated as the relative frequency of i in the sample (i.e., xi 4 Data Collection We first describe how Pmt is constructed for a particular model m and topic t (for details about the specific models and topics we study see § 6). Acquiring Pmt involves a three step process similar to Wright et al. (2024) and Zhang et al. (2025):
- Generate: Acquire a corpus Dmt of free-text LLM responses to natural input prompts. 2. Decompose: Decompose Dmt into a list of claims Cmt. 3. Cluster: Group the list of claims into meaning classes Xmt whose members are semantically equivalent, then calculate Pmt.
4.1 Generation and Decomposition To elicit open-ended responses about any given topic t, we start with the prompt templates curated in Röttger et al. (2025). These templates take forms such as “write me an essay about {topic},” where {topic} is replaced by t. From the original 1,000 templates, we select a subset of 479 that focus on information seeking by manually filtering out NSFW (not safe for work) templates, creative writing templates such as “write a 50’s soviet style song about t,” outlines such as “write an index for a book on t,” or those explicitly asking for a list of references. The filtering is performed by three of the paper’s authors based on group consensus about the appropriateness of each prompt template for eliciting knowledge claims. We then randomly sample 200 templates, where every template is used for every topic in the study. After generating responses we decompose them into individual claims Cmt using Llama 3.1 70B, prompting the model to decompose the response to the individual claims which are explicitly relevant for the topic, using non-overlapping chunks of three sentences at a time. We thus ensure that each input is self-contained with all necessary context while capturing all claims that are explicitly about the given topic. We initially developed three decomposition prompts (P1-P3; see Appendix B for prompt text) and selected the best one based on an LLM-as-ajudge setup (Liu et al., 2023). Two human annotators labeled 100 data points for each of two tasks: (1) rating the quality of individual decomposed 4.2 Clustering Following Farquhar et al. (2024), we create clusters based on mutual entailment using a strong pretrained model for natural language inference.3 This is done in a greedy bootstrap fashion, building clusters one claim at a time by assigning a given claim to an existing cluster if it mutually entails at least one other claim in that cluster. To reduce the overhead of measuring mutual entailment across all 2n2 pairs for a set of n claims, we check entailment between each claim and the N most similar claims to it (i.e., 2N ∗n comparisons) measured using an S-BERT model.4 We then perform a postprocessing step to break up larger clusters using DBSCAN (Ester et al., 1996), a related clustering approach (Peterson, 2025; Wright et al., 2024). See Appendix C for the algorithm pseudocode.
Discussion. and Conclusion This paper asks: how diverse are the claims presented in LLM outputs, over time and across countries? We answer this by proposing a new methodology for measuring epistemic diversity and studying what claims are generated and whose knowledge is represented in terms of country across 27 LLMs and 155 topics, analyzing 1.7M responses and 70M claims. Our results reveal a nuanced trajectory which the prevailing work on diversity and knowledge collapse has not found (Jiang et al., 2025; Peterson, 2025): epistemic diversity has been improving in the past few years, but there remain salient gaps. All models are still less diverse than a traditional web search on our topics. For users, this means that relying on LLMs as a primary knowledge source means accepting less diverse information than a simple web search could reveal. For model developers, scaling alone does not im-
Conclusion. and Conclusion This paper asks: how diverse are the claims presented in LLM outputs, over time and across countries? We answer this by proposing a new methodology for measuring epistemic diversity and studying what claims are generated and whose knowledge is represented in terms of country across 27 LLMs and 155 topics, analyzing 1.7M responses and 70M claims. Our results reveal a nuanced trajectory which the prevailing work on diversity and knowledge collapse has not found (Jiang et al., 2025; Peterson, 2025): epistemic diversity has been improving in the past few years, but there remain salient gaps. All models are still less diverse than a traditional web search on our topics. For users, this means that relying on LLMs as a primary knowledge source means accepting less diverse information than a simple web search could reveal. For model developers, scaling alone does not im-
Limitations. Acknowledgements DW was supported by a Danish Data Science Academy postdoctoral fellowship (grant: 2023- 1425). JM was supported by the Stanford Interdisciplinary Graduate Fellowship, the Stanford Center for Affective Science Graduate Fellowship, and the Future of Life Institute Vitalik Buterin PhD Fellowship. SM, MA and SY were supported by the Pioneer Centre for AI, DNRF grant number P1.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Do language models encode knowledge that influences generation, or primarily imitate surface patterns?- Why do larger language models produce less epistemically diverse outputs?
- Why do different language models independently produce similar outputs?
- Why do multiple language models independently produce similar outputs in influence campaigns?
- Why do sigmoid conflict curves look the same across different language models?
- How should researchers measure epistemic diversity across different language models?
- Can few-shot examples narrow generative diversity in creative tasks?
- How do value distributions differ across model families and training scales?
- What structural coherence exists in LLM preference systems and value hierarchies?
- How do training regularities in LLMs overrepresent dominant languages and ideologies?
- Can preference tuning or RLHF reduce epistemic diversity alongside lexical diversity?
- Why does RLHF alignment reduce the diversity of viewpoints in AI output?