On Epistemic Diversity in Large Language Models
Large language models (LLMs) are increasingly used not only to retrieve information, but to answer questions, explain, teach, and support inquiry. In such settings, evaluation cannot be exhausted by accuracy or alignment alone. A system may give a correct answer while still narrowing users’ access to alternative valid answers, explanations, or reasoning routes. Drawing on the broader notion of epistemic diversity in philosophy and social epistemology, we formalize it in the context of LLMs as the range of valid answers, explanations, and reasoning routes that an LLM exposes to users. We argue that epistemic diversity is a useful evaluation dimension for settings where LLMs are used to support knowledge-intensive tasks. We propose a preliminary framework for conceptualizing and measuring epistemic diversity in LLMs, and operationalize it in two domains. We find that frontier LLMs often exhibit epistemic narrowness, repeatedly collapsing large valid answer spaces onto small canonical subsets. These findings suggest that LLM evaluation should move beyond accuracy-oriented paradigms and treat epistemic diversity as an important dimension of model capability.
Introduction. Large language models (LLMs) are increasingly used not only as predictive tools, but as knowledge tools for answering questions, explaining concepts, drafting arguments, and supporting inquiry (Chatterji et al., 2025). In these settings, what matters is not only whether a model can produce a correct or acceptable answer. What also matters is the kinds of answers, explanations, and reasoning the model makes available to users. A system may be accurate and yet still be epistemically narrow if, in contexts where alternatives would be useful, it repeatedly presents only a few canonical routes despite the existence of multiple valid answers. Much of the existing literature on diversity in AI asks whether different demographic, cultural, or political groups are represented or treated fairly (Hardt et al., 2016; Barocas et al., 2023; Guo & Caliskan, 2021; Wang et al., 2025). Those are important concerns.
Discussion / Conclusion. We introduced epistemic diversity as a distinct dimension of language model evaluation: the range of valid answers, explanations, examples, concepts, and reasoning strategies that a model makes available to users. Unlike group-based diversity, epistemic diversity concerns coverage over valid answer spaces or answer classes. We formalized this idea through a framework that asks what makes an answer valid, why multiple valid answers arise, and how diversity should be measured under different interaction protocols. We operationalize this framework across ten models and two datasets, finding that frontier LLMs often exhibit epistemic narrowness, even when many valid alternatives exist. In the professions domain, models repeatedly concentrate on a small set of canonical individuals; in the proofs domain, they often return the same proof strategy despite the existence of accepted alternatives. These results suggest that models do not merely answer questions, but shape which knowledge becomes salient.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What are the consequences of models training on synthetic data? Can AI-generated outputs constitute genuine knowledge or valid claims?- Why do different AI models generate similar outputs independently?
- Can AI output be genuinely novel or only at the margins?
- Why did three experts reach incompatible conclusions about the same AI system?
- Why do different language models independently produce similar outputs?
- Why do sigmoid conflict curves look the same across different language models?
- How much alignment data does a language model actually need to specialize well?
- Why does RLHF alignment reduce the diversity of viewpoints in AI output?
- Can a single AI system optimize multiple alignment dimensions simultaneously?
- Can few-shot examples narrow generative diversity in creative tasks?
- Why does AI output show diversity without multiplying actual points of view?
- How do you verify whether your context distribution satisfies covariate diversity?
- Can diverse human creativity survive if all AI systems converge on similar outputs?
- What happens to idea diversity when AI tools draw from collective knowledge?