AI can match the average person's sense of what's 'reasonable' in law, but does it notice how different groups disagree?
Do language models track demographic variation in legal reasoning norms?
This explores whether language models pick up how legal reasoning norms differ between groups of people (by age, culture, background) or only capture a single averaged view of what's reasonable.
This is asking whether models capture how legal reasoning norms differ between groups of people, not just the average view. The corpus has no study that splits model behavior by demographic group, so there's no direct answer. The nearest evidence suggests models track the middle of the crowd well and the spread around it poorly.
The strongest legal evidence is about the aggregate. Twenty-six LLMs answered twenty-five legal reasonableness questions, and their central tendencies and distributions came fairly close to human responses. Still, the answers were llm-responses-to-legal-reasonableness-questions-approximate-human-responses-in-c|often statistically different from humans. That result compares models to people as one pooled group, so it can't say whose reasonableness the models match. A related finding hints at a shared blind spot. GPT-4.5 predicted social appropriateness better than any individual human, yet the-social-norm-savant-ai-knows-your-culture-better-than-you-do-but-from-the-out|all the AI models made identical systematic errors on unwritten norms. If models err in the same places, they are probably converging on one consensus rather than reflecting many communities.
Other kinds of variation in the corpus point the same way. Models do worse on older Supreme Court cases because llms-show-era-sensitivity-in-legal-reasoning-historical-cases-perform-worse-than|training data over-represents recent cases, which leaves shallower representations of older precedent. That is variation across eras, not demographics. But the mechanism, where whoever produces more text gets better represented, would plausibly skew toward some groups' legal intuitions too. That is my inference, not something the corpus tests. At the individual level, models individualized-reasoning-styles-distinct-reasoning-trajectories-reaching-similar|fail to track how different people reason toward similar answers, leaning on surface wording rather than the person's actual strategy. When LLMs deliberate in groups, they llm-groups-reproduce-the-human-assembly-bonus-asymmetry-in-aggregate-while-confo|converge earlier and surface less unique information than human groups, so they tend to smooth over disagreement.
Two structural notes explain why this might happen. Token prediction pushes models token-generation-is-a-smooth-probabilistic-flow-not-a-turbulent-exploration-of-r|toward the training distribution rather than toward counterpositions. A minority legal intuition is exactly the kind of counterposition that gets smoothed away. And norms are made by communities. AI can ai-can-predict-social-norms-with-superhuman-accuracy-but-cannot-participate-in-t|predict norms with superhuman accuracy yet cannot participate in creating them. A model can echo how a community reasons about legal reasonableness without registering that other communities reason differently.
The best guess from this evidence is that models track the average human view and the majority-heavy, recent-heavy slice of the text they were trained on. Whether they track the differences between groups remains untested here. A direct test would compare model answers on those reasonableness questions against human subgroups separately, then check whether model errors line up with one group's views.
Sources 7 notes
Twenty-six LLMs matched human central tendencies and distributions fairly closely on twenty-five legal reasonableness questions, with no wildly divergent means or medians, though responses were often statistically different from humans.
GPT-4.5 outperformed every individual human at judging social appropriateness across 555 scenarios, challenging the theory that embodied cultural experience is necessary. However, all AI models share identical systematic errors on unwritten norms.
Supreme Court overruling benchmark (236 pairs) reveals era sensitivity: models perform worse on historical cases than modern ones. Root cause is training corpus over-representation of recent cases, creating shallower representations of older precedent.
LLMs struggle to anchor reasoning in temporal gameplay and adapt to evolving strategies. GPT-4o relies on surface lexical cues while DeepSeek-R1 shows early promise, but dynamic style adaptation remains largely insufficient across all models tested.
LLM groups reproduce the human assembly-bonus asymmetry where discussion helps average members more than top performers, but achieve this through greater conformity, earlier convergence, and less unique information surfacing than human groups.
Show all 7 sources
Token prediction trains models to continue toward the training distribution, not to explore logically related counterpositions. This smoothness in process produces smooth claims that multiply without generating new perspectives.
GPT-4.5 outperforms all individual humans at predicting social appropriateness, yet structurally cannot enter the community processes that establish and validate norms. This reveals a critical gap between pattern-matching and authentic participation in knowledge-making.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- AI Models Exceed Individual Human Accuracy in Predicting Everyday Social Norms
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- LLM Strategic Reasoning: Agentic Study through Behavioral Game Theory
- Humans learn to prefer trustworthy AI over human partners
- SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
- The Goldilocks of Pragmatic Understanding: Fine-Tuning Strategy Matters for Implicature Resolution by LLMs
- From Process Loss to Assembly Bonus: Human-Grounded Diagnosis of Multi-Agent LLM Collaboration