Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments?

Paper · arXiv 2609.06769 · Published September 6, 2026
LLM Evaluations and Benchmarks

As people increasingly rely on artificial intelligence (AI) for guidance in their own lives, scholars, lawyers, and even judges have begun to consider the role of AI in legal decision-making. As “silicon sampling”—the use of generative AI models in social science research—is now impacting academia, “silicon jurors” could make an appearance in courtrooms. This study joins an emerging line of research on generative AI models’ ability to simulate human legal judgments. In particular, we study how large language model (LLM)-powered chatbots respond to series of questions about legal reasonableness. When the law needs to judge the appropriateness of a behavior, it most often asks whether the behavior was “reasonable.” Yet despite the ubiquity of reasonableness judgments, they are the site of constant vexation for lawyers, judges, and lay people. Reasonableness seems inherently vague and unpredictable, since it relies on variable context and implicit conceptual schemas. Moreover, many scholars caution that reasonableness judgments may vary along demographic lines. We compare the answers of human participants to those of twenty-six LLMs across twenty-five different legally relevant reasonableness judgments.

Introduction. The popularity and fluency of generative AI chatbots like ChatGPT and Claude has led people to increasingly use these tools to answer a variety of different questions—about health, careers, relationships, and beyond. Given the perceived helpfulness and convenience of using chatbots in these domains, scholars, journalists, and even federal judges have contemplated the possibility that AI could aid in various legal decision-making tasks and even, potentially, replace juries. Just as “silicon sampling” is being pushed in the social sciences (Argyle et al. 2023; Dillion et al. 2023; Bisbee et al. 2024), “silicon jurors” may soon begin to impact the law (Grimm, Grossman, and Coglianese 2024). In this paper, we contribute to the emerging research on AI’s ability to simulate human legal judgments (Posner and Saran 2025). To do so, we assess how AI chatbots respond to a series of questions about legal “reasonableness.” When the law seeks to regulate the behavior of individuals or parties, the standard it most often reaches for is “reasonableness.”

Discussion / Conclusion. and Implications responses that are similar to human responses. Although the models often provided answers that were statistically different from humans, our total impression is one of overall coherence. When looking across the range of questions in our survey, the models’ responses did a strong job of approximating the human responses–even for an inherently vague legal standard. For no question is the mean or median response from the models wildly divergent from the human response. Simple visual analysis of the violin plots in Figure 1 indicates that both the central tendencies of the models and their overall distribution of responses tend to match the human reasonableness responses. We stress that the answers to legal reasonableness judgment questions are unlikely to exist in LLMs’ latent training data in the way that the answer to a legal question like “What is the minimum age for a U.S. Senator?” will be.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can language model hallucination be prevented or only managed? Why can't humans reliably detect AI-generated text despite measurable linguistic signatures? Can AI-generated outputs constitute genuine knowledge or valid claims? How do professional roles and expertise transform with AI-generated content? Does AI fluency substitute for verifiable accuracy in human judgment? Does tokenized intelligence retain genuine value through exchange-based systems? How does AI-generated content transformation affect public discourse quality? Why do readers trust citations and complexity regardless of accuracy? Why should disagreement be treated as signal in collaborative reasoning? How do neural networks separate factual knowledge from reasoning abilities? Can prompting strategies overcome LLM biases without model fine-tuning? How should human oversight be integrated with autonomous AI systems?