On Epistemic Diversity in Large Language Models

Paper · arXiv 2609.04835 · Published September 4, 2026
LLM Evaluations and Benchmarks

Large language models (LLMs) are increasingly used not only to retrieve information, but to answer questions, explain, teach, and support inquiry. In such settings, evaluation cannot be exhausted by accuracy or alignment alone. A system may give a correct answer while still narrowing users’ access to alternative valid answers, explanations, or reasoning routes. Drawing on the broader notion of epistemic diversity in philosophy and social epistemology, we formalize it in the context of LLMs as the range of valid answers, explanations, and reasoning routes that an LLM exposes to users. We argue that epistemic diversity is a useful evaluation dimension for settings where LLMs are used to support knowledge-intensive tasks. We propose a preliminary framework for conceptualizing and measuring epistemic diversity in LLMs, and operationalize it in two domains. We find that frontier LLMs often exhibit epistemic narrowness, repeatedly collapsing large valid answer spaces onto small canonical subsets. These findings suggest that LLM evaluation should move beyond accuracy-oriented paradigms and treat epistemic diversity as an important dimension of model capability.

Introduction. Large language models (LLMs) are increasingly used not only as predictive tools, but as knowledge tools for answering questions, explaining concepts, drafting arguments, and supporting inquiry (Chatterji et al., 2025). In these settings, what matters is not only whether a model can produce a correct or acceptable answer. What also matters is the kinds of answers, explanations, and reasoning the model makes available to users. A system may be accurate and yet still be epistemically narrow if, in contexts where alternatives would be useful, it repeatedly presents only a few canonical routes despite the existence of multiple valid answers. Much of the existing literature on diversity in AI asks whether different demographic, cultural, or political groups are represented or treated fairly (Hardt et al., 2016; Barocas et al., 2023; Guo & Caliskan, 2021; Wang et al., 2025). Those are important concerns.

Discussion / Conclusion. We introduced epistemic diversity as a distinct dimension of language model evaluation: the range of valid answers, explanations, examples, concepts, and reasoning strategies that a model makes available to users. Unlike group-based diversity, epistemic diversity concerns coverage over valid answer spaces or answer classes. We formalized this idea through a framework that asks what makes an answer valid, why multiple valid answers arise, and how diversity should be measured under different interaction protocols. We operationalize this framework across ten models and two datasets, finding that frontier LLMs often exhibit epistemic narrowness, even when many valid alternatives exist. In the professions domain, models repeatedly concentrate on a small set of canonical individuals; in the proofs domain, they often return the same proof strategy despite the existence of accepted alternatives. These results suggest that models do not merely answer questions, but shape which knowledge becomes salient.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

What are the consequences of models training on synthetic data? Can AI-generated outputs constitute genuine knowledge or valid claims? Do language models learn genuine linguistic structure or just surface patterns? How does AI-generated content transformation affect public discourse quality? What makes AI persuasion effective and how can we counter it? How can AI alignment serve diverse human preferences at scale? How can language models sustain linguistic synchrony and intersubjectivity during dialogue? When does optimizing for quality undermine the value of diversity? Why does reinforcement learning suppress output diversity compared to supervised fine-tuning? Does alignment training create blind spots in detecting genuine safety threats? How do language models inherit human biases from training data? How do multi-agent systems achieve genuine cooperation and reasoning? What determines success in training models on multiple tasks? What factors beyond surface content determine how readers extract meaning differently?