Does GPT-4 actually improve the quality of legal analysis?
A randomized trial of law students using GPT-4 for legal tasks found large speed gains but small quality improvements. The finding raises questions about whether AI assistance genuinely enhances analytical capability or mainly accelerates output.
Choi, Monahan and Schwarcz report what they call the first randomized controlled trial of AI assistance on human legal analysis. Law school students were assigned realistic legal tasks with or without GPT-4, their time was tracked, and their output was blind-graded. The headline result splits speed from quality. Access to GPT-4 "only slightly and inconsistently improved the quality of participants' legal analysis but induced large and consistent increases in speed." Where quality did improve, the gain was uneven: "the lowest-skilled participants saw the largest improvements." Time saved was "roughly the same amount of time regardless of their baseline speed," and follow-up surveys found participants more satisfied and able to guess which tasks GPT-4 helped most.
The authors draw two sets of implications. Descriptively, AI can raise productivity and satisfaction and can be used selectively where it helps. Because the tool narrowed performance differences, they argue it "may also promote equality in a famously unequal profession." Normatively, they say law schools, lawyers, judges and clients should "thoughtfully embrace AI tools and plan for a future in which they will become widespread." The equality argument rests on the quality result, which is the thinnest part of the findings. The speed gain is large and consistent; the quality gain is small, and it is largest for the lowest-skilled students. A roughly constant absolute time saving also means slower students recover a larger share of their time, though the excerpt does not put it that way.
Set against the nearest notes, this trial reads differently from How often do legal AI tools actually hallucinate citations?. That note measures error rates of closed vendor products; this trial measures blind-graded analysis from a general model used by students, and finds the quality effect small. The closest parallel is Can AI narrow the education performance gap?, another randomized design in which AI narrows a gap between people of different skill, though on a different outcome and population. The sharpest tension is with Does AI turn freelance work into validation instead of creation?. The trial measures what happens within one session of work; the freelancer argument concerns what happens to skill over time. Both can be true at once, and the trial does not decide between them.
The excerpt does not establish several things. It is an abstract-length passage, so it gives no sample size, task list, effect sizes or model version beyond "GPT-4", and it studies students, not practicing lawyers. It says nothing about whether the quality gains persist, whether students who lean on AI still build the skills the tasks train, or whether the speed gain carries over to billable work. At the strength the evidence allows, the trial supports one claim: in this setting GPT-4 made legal analysis faster, and its quality benefit was small and largest for weaker students. The profession-wide equality claim is the authors' extrapolation from that result and needs evidence from practice before it can be treated as established.
Inquiring lines that read this note 3
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What are the real-world consequences of AI citation hallucinations? How can AI systems reliably guide voters without introducing political bias?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can AI narrow the education performance gap?
Does generative AI help lower-education people catch up to higher-education people on complex tasks? This matters because AI's impact on inequality depends on whether it democratizes skills or widens existing gaps.
another randomized test where AI narrows a skill-linked gap, on a different outcome.
-
How often do legal AI tools actually hallucinate citations?
Legal vendors claim their AI research tools eliminate hallucinations, but do they? This preregistered study measures hallucination rates in leading commercial legal-research systems to test those marketing claims.
contrasts vendor error rates with a trial of a general model's blind-graded legal analysis.
-
Does AI turn freelance work into validation instead of creation?
Does shifting freelancers from producing original work to validating AI output undermine their ability to build skills through paid practice? This matters because freelancers rely on client work as their primary learning mechanism.
tension: one-session speed gains versus skill formation over time, which this trial does not measure.
-
Can language models judge legal reasonableness like humans do?
Do LLMs produce judgments on vague legal standards that match human responses in both central tendency and distribution? This matters for understanding whether models can perform genuine legal reasoning rather than pattern matching.
also tests AI in legal judgment, but against human response distributions rather than trial performance.
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Lawyering in the Age of Artificial Intelligence
- How Well Can AI Do Strategy? Empirical Benchmarking Using Strategy Simulations
- Can Large Language Models Make the Grade? An Empirical Study Evaluating LLMs Ability to Mark Short Answer Questions in K-12 Education
- GPT-4 as a Homework Tutor can Improve Student Engagement and Learning Outcomes
- Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools
- How AI Impacts Skill Formation
- Evaluating Large Language Models in Theory of Mind Tasks
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Original note title
GPT-4 made law students much faster at legal analysis but only slightly and inconsistently better — the lowest-skilled gained most