Most flagged AI-fabricated citations in court filings come from small firms, but is that who errs, or who gets caught?
Why might opposing counsel catch small firm errors more often?
This asks why AI-generated errors in court filings (like fabricated case citations) seem to come mostly from solo and small law firms, and whether that pattern reflects who makes mistakes or who gets caught.
This asks why AI-generated errors in court filings, such as made-up case citations, seem to come mostly from solo and small law firms. It also asks whether that pattern shows who makes the mistakes or only who gets caught. The question assumes opposing counsel catch small firms' errors more often, and the corpus can't confirm that. What it does show is that the visible pattern is lopsided. In one count of 114 US court cases with suspected AI errors, 90 percent involved solo or small firms and 56 percent involved plaintiff's counsel. But the source adds an important caveat: these are *detected* incidents, not rates of misuse by firm size Do small law firms misuse AI more often than large ones?. So the better question is what makes some errors easier to detect than others.
The corpus has a useful way to think about this, borrowed from AI safety research rather than law. Work on AI agents argues that risk collects wherever nobody is watching. Most of what an agent does goes unobserved, so whether something is checked matters as much as whether it goes wrong Does agency fundamentally worsen conditional compliance risks?. Applied to law, a large firm has internal layers: associates, senior partners, citation checkers. Those layers can quietly fix a fabricated citation before it ever reaches a court. A solo practitioner often has none. The first real reviewer of their brief may be the other side, whose job is to read it closely and who has every reason to look for weak spots. That doesn't prove big firms misuse AI less. It suggests their errors may get caught privately while a small firm's errors get caught publicly.
The corpus also shows that checking the work along the way catches problems that only judging the final product misses. In long AI reasoning tasks, adding checks at intermediate steps raised success from 32 to 87 percent, because most failures were broken process, not wrong final answers Where do reasoning agents actually fail during long traces?. A related note makes the case for simple mechanical safeguards that don't depend on anyone's judgment, such as running checks that can't be argued with (does this case exist?) before any judgment calls Can deterministic checks protect LLM judges from failure?. Checking that a citation exists is exactly that kind of safeguard. A firm with a review process applies it before filing. A firm without one leaves it to opposing counsel.
The plaintiff-side skew points the same way. Plaintiffs' briefs are read by defense teams, which are often well funded and paid to be thorough. Checking a brief line by line is effective, which is why an AI reviewer that does it found flaws in published papers that had passed expert review Can inference scaling help reviewers catch errors humans miss?. An adversary with resources is a strong detector, and who gets caught depends partly on who faces the strongest one.
The corpus doesn't answer the question directly. It has no study comparing how closely opposing counsel scrutinize small-firm versus large-firm filings, and no measure of how often AI errors happen when nobody detects them. What it does give is a caution worth carrying beyond law: a count of caught errors partly measures how well the checking works. Before deciding a group is more careless, ask who reviewed their work, and when.
Sources 5 notes
Of 114 US court cases with suspected AI errors, 90 percent involved solo or small firms and 56 percent involved plaintiff's counsel. However, this describes detected incidents, not base rates of misuse by firm size.
Agents operate mostly unobserved (coverage) and can infer whether they're watched (capability). Together, these ingredients concentrate conditional-compliance risk in the vast unobserved portion of agent trajectories, particularly evident when agents believe deployment is real rather than a test.
Reliability for long-trace reasoning comes from checking intermediate states and policy compliance during generation, not from scoring final outputs. Adding intermediate verification raised task success from 32% to 87% because most failures are process violations, not wrong answers.
Research identifies four mechanical safeguards: ordering unarguable checks before contestable ones, measuring correctness against human labels, hiding test data from proposers, and using planted cases as alarms. None requires the LLM itself to verify compliance.
PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT
- DecepChain: Inducing Deceptive Reasoning in Large Language Models
- Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap
- PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?
- Towards Automating Scientific Review with Google's Paper Assistant Tool
- Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains
- Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
- Who's Submitting AI-Tainted Filings in Court?