Can AI-assisted reports pass quality checks with fabricated citations?
A Deloitte government report contained over a dozen invented references and fake quotes that escaped internal review. The question is whether AI-generated content can systematically bypass citation verification in high-stakes professional work.
Deloitte will partially refund Australia's Department of Employment and Workplace Relations (DEWR) $440,000 after a report it was paid to write turned out to contain "inaccurate footnotes and made up references and quotes." University of Sydney academic Chris Rudge found "over a dozen fabricated points" attributed to colleagues whose work he knew personally: "I was in no doubt, when I read the names of the works, that they were fake because I work in the area and I know my colleagues' work." Among the fabrications: a quote attributed to "Justice Davis" in a 2019 Federal Court case (misspelling Justice Davies) citing paragraphs 25 and 26 of consent orders that contain no such paragraphs, and two cited papers — one in the UNSW Law Journal — that do not exist.
Deloitte confirmed it used "a generative artificial intelligence (AI) large language model (Azure OpenAI GPT-4o) based tool chain licensed by DEWR and hosted on DEWR's Azure tenancy" in producing the report, but has not confirmed the fabrications were caused by that tool chain; a spokesperson said only that "the matter has been resolved directly with the client." DEWR's own statement drew a sharp line between the errors and the work's conclusions: "The substance of the independent review is retained, and there are no changes to the recommendations." Rudge's own read inverts the usual framing of AI error as the worse outcome: "I think it's good if it's AI, because to think of a person doing that is almost worse."
The fabricated-citation mechanism matches Do small law firms misuse AI more often than large ones?, where hallucinated case law slipped into court filings past the filer's own check. This case extends that pattern from solo and small-firm litigation to a named Big Four consultancy delivering a six-figure government contract — the same failure mode (invented sources that read as real) surfacing in higher-stakes, better-resourced professional work. It also sits against Does disclosing AI use damage how trustworthy you seem?: Deloitte disclosed the tool chain only after outside detection forced the question, and even then declined to say whether the AI caused the errors, consistent with a disclosure cost that gives firms reason to confirm as little as possible.
The excerpt is one detected incident, not a rate: it does not say how many other DEWR-funded reports used the same tool chain without errors being caught, nor whether Deloitte's internal review process checked citations before publication and failed, or skipped that check entirely. It also does not establish that AI caused the fabrications — Deloitte never confirmed this, and the report's content changes (two silent revisions before the refund) suggest the firm treated citation accuracy as separable from substance. The incident supports, at the strength of a single case, that a known-reputable firm's AI-assisted output can carry fabricated citations past its own quality control into a paid government deliverable; it does not support a claim about how often that happens.
Inquiring lines that read this note 4
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What are the real-world consequences of AI citation hallucinations? How do hallucinated citations emerge in AI scholarly output? What governance mechanisms can effectively constrain widely deployed AI systems? How do educators verify student capability when AI can produce indistinguishable work?Related concepts in this collection 2
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do small law firms misuse AI more often than large ones?
A database of 114 court cases with AI-tainted filings shows 90 percent involved small or solo firms. But does this reflect higher misuse rates, or simply better detection of errors in smaller practices?
same fabricated-citation failure mode, here surfacing in a Big Four consultancy's paid government report rather than small-firm litigation
-
Does disclosing AI use damage how trustworthy you seem?
When people learn you used AI to create work, do they trust you less? Schilke and Reimann tested this across 13 experiments with over 5,000 participants to understand whether transparency about AI reliance backfires.
Deloitte disclosed its AI tool chain only once forced to, and declined to confirm AI caused the errors, consistent with a disclosure cost
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Deloitte to refund government after using AI in $440k report
- Your AI Strategy Advisor Is Giving Everyone the Same Advice
- Deloitte refunds Australian government for AI-flawed report
- Tortured phrases: A dubious writing style emerging in science. Evidence of critical issues affecting established journals
- Pangram Predicts 21% of ICLR Reviews are AI-Generated
- Stranded Credentials: Keeping Online Reputation Systems Informative in the AI Era
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
Original note title
Deloitte refunds the Australian government part of a $440,000 fee after AI fabricated references and quotes in its report