SYNTHESIS NOTE
Topics›AI at Work›this note

Can AI quality control catch fabricated citations in professional reports?

This explores whether existing review processes at major firms can catch AI-generated errors like fake quotes and nonexistent citations before client delivery. It matters because firms claim to maintain quality oversight while adopting AI tools.

Synthesis note · 2026-10-09 · sourced from AI at Work

CFO Dive reports that Deloitte Australia refunded "over 97,000 Australian dollars ($63,000 USD)" to Australia's Department of Employment and Workplace Relations (DEWR) after a paid assurance report it delivered was found to contain "artificial intelligence-generated errors." The refund covers less than a quarter of the roughly A$440,000 contract, executed in December 2024, for a review of the IT system the department uses to automate welfare penalties. According to the Associated Press, as cited by CFO Dive, the original report contained "a fabricated quote from a federal court judgment and references to nonexistent academic research papers."

DEWR's spokesperson said Deloitte "conducted the independent assurance review and has confirmed some footnotes and references were incorrect," while adding that "the substance" of the review had been retained — drawing a line between the report's citations and footnotes, where the AI-generated failure occurred, and its substantive findings, which the client maintains survived. Hofstra accounting professor Jack Castonguay frames the failure as a training and quality-control gap rather than a one-off glitch, telling CFO Dive that firms "need to not only train staff on how to use AI effectively, but they must also train them on ethical use and maintaining quality control," and that AI output "must be reviewed as if it were prepared by an intern or new hire."

The incident is a concrete instance of the failure mode described in Can organizations lose scrutiny capacity while keeping oversight forms?: Deloitte presumably had some review process in place for a government-facing assurance report, yet fabricated citations reached the delivered, paid document, meaning the scrutiny step that should have caught invented quotes and papers did not function as intended. It also sits in tension with How close are frontier AI models to expert work quality?: that benchmark finds frontier models approaching expert quality on judged work tasks, but here a live, paid professional deliverable from a major firm shipped with fabricated legal and academic citations, pointing to a gap between benchmarked capability and quality control under real client conditions. The public exposure and partial refund also give the Does disclosing AI use damage how trustworthy you seem? dynamic a firm-level, involuntary analog: credibility costs that Schilke and Reimann find at the level of an individual choosing to disclose AI use appear here as a reputational and contractual cost imposed on a firm after AI-generated errors were discovered by others.

The excerpt does not establish how the errors were produced — which AI tool was used, whether its use was disclosed to DEWR in advance, or whether the review failure traces to time pressure, understaffing, or a specific breakdown in Deloitte's checking process. It also gives no basis for estimating how common this kind of failure is across Deloitte's or other Big Four firms' AI-assisted deliverables, since this is one disclosed case that became public rather than a sampled or audited rate. The narrow, supportable implication is that fabricated AI output can reach a paid, government-facing professional deliverable despite an "independent assurance review," and that when caught, it carries a direct financial consequence; whether review processes industry-wide are adequate to catch such errors before delivery is not something this source speaks to.

Inquiring lines that read this note 5

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What are the real-world consequences of AI citation hallucinations? How do hallucinated citations emerge in AI scholarly output?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 137 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Deloitte refunds the Australian government after its AI-assisted report fabricated a court quote and nonexistent research papers