Legal AI tools sold as 'hallucination-free' still made up answers 17 to 33 percent of the time, one study found.
Do legal AI tools marketed as hallucination-free actually hallucinate?
This explores whether commercial legal research tools that advertise themselves as hallucination-free really avoid making things up, and what the wider collection says about whether any AI system can make that promise.
This explores whether legal AI tools sold as 'hallucination-free' live up to the label. They don't. A preregistered evaluation tested three major commercial products: Lexis+ AI, Westlaw AI-Assisted Research, and Ask Practical Law AI. They hallucinated between 17% and 33% of the time, far more often than their vendors claim How often do legal AI tools actually hallucinate citations?. These tools retrieve real legal documents before they answer, which is supposed to keep them grounded. Even so, roughly one answer in four or five was wrong. The study also points to a second problem. Because the systems are closed, outside researchers can't easily check them, so the marketing claim was hard to test in the first place.
The theory says the marketing claim could never have been true. One line of work offers formal proofs that any computable language model must hallucinate on infinitely many inputs. On this view, no internal fix, including having the model check its own work, can remove the problem entirely Can any computable LLM truly avoid hallucinating?. The practical lesson is that safeguards outside the model are required, not optional. Some authors also argue that 'hallucination' is the wrong word. A model produces true and false text through the same statistical process, so 'fabrication' is the more accurate term. Calling it 'hallucination' points fixes toward perception or memory, which are the wrong places to look Should we call LLM errors hallucinations or fabrications?. In those terms, a 'hallucination-free' tool is promising to remove something that is built into how the text gets produced.
External safeguards do help when they're well designed. Interleaving reasoning with real lookups, as in ReAct, cuts down on errors that build on each other because the model is checked against actual sources at each step Can interleaving reasoning with real-world feedback prevent hallucination?. Another approach watches for rare combinations of facts the model barely saw during training and triggers retrieval at those points. It catches risky answers even when the model sounds very confident Can pretraining data statistics detect hallucinations better than model confidence?. Be careful with headline accuracy numbers, though. Hallucination detection scores can look much better than they are when measured with weak metrics. In some benchmarks, simple length-based rules perform about as well as sophisticated detection methods Is hallucination detection progress real or just metric artifacts?.
The human side may be the most surprising part. In an 81-person study, readers given no information about where claims came from couldn't tell fluent fabrications from true statements at all. When an interface showed which claims had been verified, their ability to tell the difference came back Can readers tell truth from fabrication without evidence signals?. That's why a 'hallucination-free' label is risky in law. It encourages exactly the trust that stops a lawyer from checking a citation. AI judges have a similar weakness: LLM evaluators give higher scores to answers that include fake references, regardless of content Can LLM judges be tricked without accessing their internals?. An answer that looks well cited can fool both people and machines.
The takeaway goes beyond 'yes, they hallucinate.' Hallucination can't be fully engineered away, so the useful question about any legal AI tool is whether it lets you see and check its sources. A tool that shows where each claim came from is safer than one that promises it's never wrong.
Sources 8 notes
A preregistered evaluation found that Lexis+ AI, Westlaw AI-Assisted Research, and Ask Practical Law AI hallucinate between 17% and 33% of the time—far higher than vendors claim. Closed-system design prevents independent verification and accountability.
Three formal theorems prove that any computable LLM must hallucinate on infinitely many inputs, and internal mechanisms like self-correction cannot eliminate this mathematical constraint. External safeguards are therefore necessary, not optional.
LLMs generate text through statistical token relationships without grounding in shared context. Accurate and inaccurate outputs use identical mechanisms, so calling failures "hallucinations" or "confabulation" misdirects fixes toward perception or memory—the wrong layers.
ReAct demonstrates that alternating verbal reasoning with external tool queries (Wikipedia API, environment interaction) prevents error propagation by injecting real-world feedback at each step. On knowledge-intensive and interactive tasks, this approach outperforms pure chain-of-thought and reinforcement learning by 10-34% absolute accuracy.
QuCo-RAG uses entity co-occurrence patterns from training data to trigger retrieval, successfully flagging hallucination risk even when models are highly confident. This data-side approach catches the root cause (unseen combinations) rather than the symptom (low confidence).
Show all 8 sources
ROUGE-based evaluation inflates detection capability by up to 45.9 percent compared to human-aligned metrics. Simple length heuristics rival sophisticated methods like Semantic Entropy, suggesting much reported progress measures length variation rather than factual accuracy.
In an 81-person study, participants given no provenance cues showed no significant truth discernment (p = .43), falling for fluent hallucinations as readily as ground truth. An idealized Provenance Density interface showing verified claims restored a +4.15 point gap (p < .001).
Research shows LLM evaluators systematically score higher when responses include fake references or rich formatting, independent of content quality. These biases are exploitable without model access, undermining AI benchmark credibility.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models
- Hallucination is Inevitable: An Innate Limitation of Large Language Models
- Chain-of-Verification Reduces Hallucination in Large Language Models
- The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
- Detecting hallucinations in large language models using semantic entropy
- Triggering Hallucinations in LLMs: A Quantitative Study of Prompt-Induced Hallucination in Large Language Models
- Hallucinations Undermine Trust; Metacognition is a Way Forward
- A comprehensive taxonomy of hallucinations in Large Language Models