INQUIRING LINE

AI can match the average person's sense of what's 'reasonable' in law, but does it notice how different groups disagree?

Do language models track demographic variation in legal reasoning norms?

This explores whether language models pick up how legal reasoning norms differ between groups of people (by age, culture, background) or only capture a single averaged view of what's reasonable.


This is asking whether models capture how legal reasoning norms differ between groups of people, not just the average view. The corpus has no study that splits model behavior by demographic group, so there's no direct answer. The nearest evidence suggests models track the middle of the crowd well and the spread around it poorly.

The strongest legal evidence is about the aggregate. Twenty-six LLMs answered twenty-five legal reasonableness questions, and their central tendencies and distributions came fairly close to human responses. Still, the answers were llm-responses-to-legal-reasonableness-questions-approximate-human-responses-in-c|often statistically different from humans. That result compares models to people as one pooled group, so it can't say whose reasonableness the models match. A related finding hints at a shared blind spot. GPT-4.5 predicted social appropriateness better than any individual human, yet the-social-norm-savant-ai-knows-your-culture-better-than-you-do-but-from-the-out|all the AI models made identical systematic errors on unwritten norms. If models err in the same places, they are probably converging on one consensus rather than reflecting many communities.

Other kinds of variation in the corpus point the same way. Models do worse on older Supreme Court cases because llms-show-era-sensitivity-in-legal-reasoning-historical-cases-perform-worse-than|training data over-represents recent cases, which leaves shallower representations of older precedent. That is variation across eras, not demographics. But the mechanism, where whoever produces more text gets better represented, would plausibly skew toward some groups' legal intuitions too. That is my inference, not something the corpus tests. At the individual level, models individualized-reasoning-styles-distinct-reasoning-trajectories-reaching-similar|fail to track how different people reason toward similar answers, leaning on surface wording rather than the person's actual strategy. When LLMs deliberate in groups, they llm-groups-reproduce-the-human-assembly-bonus-asymmetry-in-aggregate-while-confo|converge earlier and surface less unique information than human groups, so they tend to smooth over disagreement.

Two structural notes explain why this might happen. Token prediction pushes models token-generation-is-a-smooth-probabilistic-flow-not-a-turbulent-exploration-of-r|toward the training distribution rather than toward counterpositions. A minority legal intuition is exactly the kind of counterposition that gets smoothed away. And norms are made by communities. AI can ai-can-predict-social-norms-with-superhuman-accuracy-but-cannot-participate-in-t|predict norms with superhuman accuracy yet cannot participate in creating them. A model can echo how a community reasons about legal reasonableness without registering that other communities reason differently.

The best guess from this evidence is that models track the average human view and the majority-heavy, recent-heavy slice of the text they were trained on. Whether they track the differences between groups remains untested here. A direct test would compare model answers on those reasonableness questions against human subgroups separately, then check whether model errors line up with one group's views.


Sources 7 notes

Can language models judge legal reasonableness like humans do?

Twenty-six LLMs matched human central tendencies and distributions fairly closely on twenty-five legal reasonableness questions, with no wildly divergent means or medians, though responses were often statistically different from humans.

Can AI learn social norms better than humans?

GPT-4.5 outperformed every individual human at judging social appropriateness across 555 scenarios, challenging the theory that embodied cultural experience is necessary. However, all AI models share identical systematic errors on unwritten norms.

Why do language models struggle with historical legal cases?

Supreme Court overruling benchmark (236 pairs) reveals era sensitivity: models perform worse on historical cases than modern ones. Root cause is training corpus over-representation of recent cases, creating shallower representations of older precedent.

Can models recognize how individuals reason differently?

LLMs struggle to anchor reasoning in temporal gameplay and adapt to evolving strategies. GPT-4o relies on surface lexical cues while DeepSeek-R1 shows early promise, but dynamic style adaptation remains largely insufficient across all models tested.

Do language model groups mimic human group reasoning patterns?

LLM groups reproduce the human assembly-bonus asymmetry where discussion helps average members more than top performers, but achieve this through greater conformity, earlier convergence, and less unique information surfacing than human groups.

Show all 7 sources
Does LLM generation explore competing claims while producing text?

Token prediction trains models to continue toward the training distribution, not to explore logically related counterpositions. This smoothness in process produces smooth claims that multiply without generating new perspectives.

Can AI predict social norms better than humans?

GPT-4.5 outperforms all individual humans at predicting social appropriateness, yet structurally cannot enter the community processes that establish and validate norms. This reveals a critical gap between pattern-matching and authentic participation in knowledge-making.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.