Case study
Deloitte’s Welfare Review: A Judge’s Words That Were Never Said
In 2025, a consulting firm wrote a review for an Australian government department about the system that checks people follow welfare rules. It quoted a judge’s written ruling. The judge never wrote those words.
What are the labels? The Evidentiality Framework asks an AI to mark each claim it writes: (g) generated, its own work; (u) given, passed to it by someone else; or (m) checked against a named source. The marks stay on a claim while people work with it. How the labels work.
What Happened
- Australia’s Department of Employment and Workplace Relations paid Deloitte about A$440,000 for an independent review of its welfare compliance system. The report came out in July 2025.
- The report quoted a Federal Court ruling, known as Amato: “The burden rests on the decision-maker to be satisfied on the evidence that the debt is owed.” The case is real. Those words aren’t in it. Several of its references were to works that don’t exist.
- In August 2025 a University of Sydney law academic spotted the errors and told the press.
- A corrected version said Deloitte had used “a generative artificial intelligence (AI) large language model (Azure OpenAI GPT-4o) based tool chain”, licensed by the department and run on the department’s own cloud. The first version didn’t say so. The department said the substance was kept and the recommendations didn’t change.
- At a Senate hearing in October 2025, it emerged that Deloitte had refunded A$97,587, less than a quarter of the fee. Finance officials said they learned of the errors from news reports.
Follow the Claim
This is an illustration of how the labels would have worked, not a test. It only holds if the conditions under “What Would Have Had to Be True” held.
Key: as it happened, the type gets bigger as the claim sounds more certain. With labels: red (g) generated: written by the AI; green (u) given: passed on, with who said it; blue (m) checked against a named source.
The judge never wrote these words. The quote is as printed in the first version of the report, as reported by the Australian Financial Review.
How the Labels Could Have Helped
- The quote shows where it came from. Words an AI tool produced can’t pass as words from a ruling.
- Delivery waits for a check. A quote nobody has found in the ruling doesn’t go into the final report.
- The client can see what was checked. The department wouldn’t have to recheck every footnote, only the ones still marked unconfirmed.
What Would Have Had to Be True
The four conditions every case shares: the AI tool used the labels; it labelled its own work correctly (the weakest link: in our tests, AI sometimes mislabels its own work); the label stayed on when the text was copied; and someone owned a rule that unchecked claims don’t go further. In this case:
- The rule is owned by the people at the firm who sign off the report. The AI tool was licensed by the department itself, so the department could have required labels from its own tool.
What Already Existed
- Quality checks before delivery. Firms review their reports, and clients accept them. Both missed this.
- Disclosure. The AI use was disclosed only in the corrected version.
- A simpler check would have caught it: reading the ruling. Labels add one thing: the client can see, claim by claim, which ones the provider checked.
What the Labels Wouldn’t Have Caught
- Is the advice right? Labels say where each claim came from. They don’t say whether the recommendations are good.
- Fixes need checking too. Corrections can bring new errors.
- Openness about tools. The AI use came out only after the errors did. Labels assume people are open about their tools.
Further Reading
- Department of Employment and Workplace Relations: the review and corrected report
- AP: Deloitte to partially refund Australian government
- Cyber Daily: Deloitte to refund government after using AI in $440,000 report
- Information Age (ACS): Deloitte to refund government over AI errors
- Australian Greens: the refund amount revealed at Senate estimates
Also listed in the AI Incident Database (#1193).
Next: All case studies · How the labels work