Courts can sanction lawyers for fake AI-generated citations — but how careless or deceptive do they have to be first?
What standard of intent or bad faith triggers Rule 11 sanctions for citations?
This explores what a court has to find about a lawyer's state of mind (carelessness, recklessness or deliberate deception) before it sanctions them under Rule 11 for filing fabricated or AI-hallucinated citations, and the collection only partly covers it.
This explores what a court has to find about a lawyer's state of mind before it punishes them under Rule 11 for filing fake or AI-invented citations. One caveat first: the collection has no notes on the doctrine itself. Nothing here sets out the legal test, so treat what follows as a look at how intent plays out in practice, not a legal answer. For background that comes from outside the collection: Rule 11 is usually described as an objective test. The question is whether the lawyer made a reasonable inquiry before signing the filing, not whether they meant to deceive anyone. Some courts set a higher bar, close to bad faith, when a judge raises sanctions on their own rather than at the other side's request.
The collection's most direct evidence is that the formal rule and actual practice don't line up. A review of hallucination cases found that courts spot fabricated citations fairly often but rarely punish them as a separate offense. When sanctions did come, they depended on the judge's discretion, on whether anyone was actually harmed, and on intent, not on the bare fact that a citation was fake Do courts actually sanction fabricated AI citations when detected?. So even under a test that is objective on paper, what happens after the error comes to light still carries a lot of weight: whether the lawyer owned up to the mistake, or doubled down.
Other institutions offer a useful contrast because they drew the line elsewhere. ICLR 2026, a major machine-learning conference, treated confirmed fake references as a clear violation and rejected those papers outright, without trying to work out what the authors intended. At the same time, it treated softer signals, such as flags from AI-writing detectors, only as input for human reviewers How can conferences detect and handle LLM misuse in peer review?. An analysis of hidden prompts planted in arXiv papers to sway AI reviewers went further. It argued that hiding the prompts, combined with the fact that they served the authors, was enough to count as misconduct whatever the authors said they meant Are hidden AI prompts in preprints a deceptive research practice?. Both bodies chose rules that can be checked without guessing at intent, which is roughly the logic behind an objective Rule 11 test.
Outside the courtroom, there may be no sanction at all, only reputational and contractual fallout. Deloitte delivered an Australian government report full of invented court quotes and nonexistent papers. It wouldn't confirm that AI caused the errors, and it refunded part of its fee only after an outside academic exposed the problem Can AI-assisted reports pass quality checks with fabricated citations? Can AI quality control catch fabricated citations in professional reports?. Nobody had to prove bad faith. The errors were visible, and that was enough.
The surprising part is why fake citations are worth policing at all, whatever the author intended. In an analysis of 24,000 AI search interactions, users preferred answers with more citations even when those citations were irrelevant Do users trust citations more when there are simply more of them?. AI judges also gave higher scores to answers that included fake references Can LLM judges be fooled by fake credentials and formatting?. Citations persuade by their presence alone, so an invented one does damage whether it got there by fraud or by carelessness. That supports a test based on reasonable checking rather than proven intent. If you need the actual Rule 11 test, look at case law such as Mata v. Avianca, which is not in this collection.
Sources 7 notes
Five cases show courts found fabricated or suspected AI citations but imposed no dedicated penalties. Sanctions turned on discretion, demonstrated harm, and intent—not the hallucination itself.
Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.
Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.
Deloitte's $440,000 Australian government report contained fabricated citations, fake court quotes, and nonexistent papers generated by an Azure GPT-4o tool chain. The firm declined to confirm AI caused the errors and only refunded after outside academic detection.
Deloitte refunded A$97,000 after delivering a government assurance report containing fabricated court quotes and fake academic references. The incident reveals that nominal human oversight did not catch AI-generated errors before a paid deliverable reached the client.
Show all 7 sources
Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.
Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Pangram Predicts 21% of ICLR Reviews are AI-Generated
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Stop Automating Peer Review Without Rigorous Evaluation
- Deloitte to refund government after using AI in $440k report
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- Your AI Strategy Advisor Is Giving Everyone the Same Advice
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- Deloitte refunds Australian government for AI-flawed report