This week Deloitte Australia refunded part of a AU$440,000 contract after a government report it delivered was found to contain fabricated academic references and a misquoted Federal Court judgment, errors consistent with unchecked AI output. The story has been reported worldwide, mostly as a cautionary tale about AI. We think that's the wrong lesson. Nobody involved is going to stop using AI, and neither should you. What failed here wasn't the technology. It was the checking.
What You Need to Know
- The report: an independent assurance review for Australia's Department of Employment and Workplace Relations, published July 2025, under a contract worth about AU$440,000.
- The errors: references to academic papers that don't exist, fabricated work attributed to real researchers, and a misquoted judgment from the Amato robodebt case, with a Justice's name misspelt.
- The discovery: not an internal review. A University of Sydney welfare academic, Dr Chris Rudge, found over a dozen fabricated citations, and journalists took it from there.
- The outcome: Deloitte disclosed that an Azure OpenAI GPT-4o toolchain was used in producing the report, republished a corrected version, and refunded the final instalment of over AU$97,000.
What Actually Failed
Look at the specific errors. Non-existent paper titles attributed to real academics. A quote that was never in the judgment. These are classic generative AI hallucinations: fluent, plausible, and checkable in minutes by anyone who tries. Dr Rudge put it perfectly: he couldn't understand how a human could invent the titles of works that don't appear on Google. A human didn't. And no human checked.
That's the entire failure, and it's worth being precise about it, because none of the following would have made headlines:
- Using AI to draft the report. The client in this case had licensed the AI tooling themselves; it ran on the department's own Azure tenancy.
- AI producing a wrong citation in a draft. Models do this. It's a known, documented property of the technology, which is exactly why the checking step exists.
What made headlines is that fabricated citations survived every review between a draft and a published government document, and that the reader who finally checked them was outside the organisation. The failure was a process that treated AI output as finished work.
You Own What You Ship
Here's the principle we hold ourselves and our clients to, and the one this story teaches: you are responsible for every output you use, however it was produced. The tool is never accountable. The signature on the report is.
That principle isn't new. Firms have always been liable for the work of juniors, subcontractors, and spreadsheets. AI doesn't change the accountability; it changes the throughput, which means the verification discipline has to scale with it. A tool that can draft in minutes what took a team weeks demands checking that's designed in, not bolted on:
- Citations get verified, mechanically. Every reference in an AI-assisted document should be resolvable to a real source before it leaves the building. This is automatable. A checker that confirms each cited work exists and each quote appears in its source would have caught every error in this report.
- Grounding beats generation. The deeper fix is not letting the model free-associate in the first place: retrieval-first systems that answer from real documents, and that say so when a source doesn't exist. A model asked to support an argument will oblige; a grounded system asked the same question can decline.
- Disclosure up front, not after discovery. The AI use became a story partly because it surfaced in the corrected version, after a researcher and a journalist forced the question. If AI materially contributes to a deliverable, say so at delivery. Disclosure costs a sentence; discovery costs the contract and the headline.
- Human review means adversarial review. A skim is not a check. Someone whose name is on the work needs to attack the claims the way Dr Rudge did: pick citations and try to find them. If your review process wouldn't catch a fabricated reference, it isn't a review process.
The Right Amount of AI Is Not Zero
The tempting response, especially in government and regulated sectors, is to ban the tools. That response fails twice: it forfeits the productivity that every serious study now attributes to well-deployed AI, and it doesn't even remove the risk, because staff use AI anyway, just invisibly and without any of the controls above.
The mature response is the one the entire industry got handed this month at someone else's expense: use AI, ground it, check it, disclose it, and own it. We've argued before that AI governance is not optional, and that governance starts with what feeds the system. This story adds the third leg: governance ends with what leaves the building, and that part is never the machine's job.