Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments?
As people increasingly rely on artificial intelligence (AI) for guidance in their own lives, scholars, lawyers, and even judges have begun to consider the role of AI in legal decision-making. As “silicon sampling”—the use of generative AI models in social science research—is now impacting academia, “silicon jurors” could make an appearance in courtrooms. This study joins an emerging line of research on generative AI models’ ability to simulate human legal judgments. In particular, we study how large language model (LLM)-powered chatbots respond to series of questions about legal reasonableness. When the law needs to judge the appropriateness of a behavior, it most often asks whether the behavior was “reasonable.” Yet despite the ubiquity of reasonableness judgments, they are the site of constant vexation for lawyers, judges, and lay people. Reasonableness seems inherently vague and unpredictable, since it relies on variable context and implicit conceptual schemas. Moreover, many scholars caution that reasonableness judgments may vary along demographic lines. We compare the answers of human participants to those of twenty-six LLMs across twenty-five different legally relevant reasonableness judgments.
Introduction. The popularity and fluency of generative AI chatbots like ChatGPT and Claude has led people to increasingly use these tools to answer a variety of different questions—about health, careers, relationships, and beyond. Given the perceived helpfulness and convenience of using chatbots in these domains, scholars, journalists, and even federal judges have contemplated the possibility that AI could aid in various legal decision-making tasks and even, potentially, replace juries. Just as “silicon sampling” is being pushed in the social sciences (Argyle et al. 2023; Dillion et al. 2023; Bisbee et al. 2024), “silicon jurors” may soon begin to impact the law (Grimm, Grossman, and Coglianese 2024). In this paper, we contribute to the emerging research on AI’s ability to simulate human legal judgments (Posner and Saran 2025). To do so, we assess how AI chatbots respond to a series of questions about legal “reasonableness.” When the law seeks to regulate the behavior of individuals or parties, the standard it most often reaches for is “reasonableness.”
Discussion / Conclusion. and Implications responses that are similar to human responses. Although the models often provided answers that were statistically different from humans, our total impression is one of overall coherence. When looking across the range of questions in our survey, the models’ responses did a strong job of approximating the human responses–even for an inherently vague legal standard. For no question is the mean or median response from the models wildly divergent from the human response. Simple visual analysis of the violin plots in Figure 1 indicates that both the central tendencies of the models and their overall distribution of responses tend to match the human reasonableness responses. We stress that the answers to legal reasonableness judgment questions are unlikely to exist in LLMs’ latent training data in the way that the answer to a legal question like “What is the minimum age for a U.S. Senator?” will be.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can language model hallucination be prevented or only managed?- Does retrieval augmented generation actually eliminate hallucinations in any domain?
- Can architectural changes reduce hallucination without external retrieval or verification?
- What happens when lawyers rely on AI citations that turn out false?
- Can social validation of expertise exclude systems that lack participatory track records?
- What does it mean that AI knowledge is structurally hearsay?
- Can AI systems produce genuinely new validity claims without community participation?
- Can diverse expert demonstrations exceed the knowledge of any single expert?
- What expertise survives in a world where AI can generate knowledge on demand?
- What role shifts occur when experts become custodians of AI knowledge?
- What role did human experts play in raising social alarms historically?
- What happens to professional expertise when judgment gets encoded into systems?
- How does AI presentation authority substitute for actual expert judgment?
- Does surface authority without earned authority create risks in expert judgment?
- What makes expert judgment depend on anticipating audience acceptability?
- What happens to expert credibility when AI-generated claims drown out specialist signals?
- Does stripping social context from knowledge claims hollow out their meaning?