Evidentiality Framework for AI
Download .md

Case study

Deloitte’s Welfare Review: A Judge’s Words That Were Never Said

In 2025, a consulting firm wrote a review for an Australian government department about the system that checks people follow welfare rules. It quoted a judge’s written ruling. The judge never wrote those words.

What are the labels? The Evidentiality Framework asks an AI to mark each claim it writes: (g) generated, its own work; (u) given, passed to it by someone else; or (m) checked against a named source. The marks stay on a claim while people work with it. How the labels work.


What Happened


Follow the Claim

This is an illustration of how the labels would have worked, not a test. It only holds if the conditions under “What Would Have Had to Be True” held.

Follow one quote: from an AI tool to a government review

Key: as it happened, the type gets bigger as the claim sounds more certain. With labels: red (g) generated: written by the AI; green (u) given: passed on, with who said it; blue (m) checked against a named source.

As it happenedWith labels
2025 · AI tool → report authors
As it happened“The burden rests on the decision-maker to be satisfied on the evidence that the debt is owed.”Words put in a judge’s mouth. The case is real; the quote isn’t.
With labels(g)The burden rests on the decision-maker to be satisfied on the evidence that the debt is owed.(/g)If the tool used the labels: marked as written by the AI.
2025 · Authors → the draft report
As it happenedWritten into the draft as a court quoteNow it reads as legal authority.
With labels(u)The burden rests on the decision-maker to be satisfied on the evidence that the debt is owed.(/u: AI tool, unconfirmed)Passed on with where it came from, and that nobody has read it in the ruling.
Before July 2025 · The check
As it happenedNobody read the rulingThe quote was never looked up.
With labels(m)The ruling contains no such words(/m: Australian Financial Review, 5 October 2025, checked 3 Oct 2026)This check could be made at the time. We name the source that confirmed it later.
July 2025 · Deloitte → the department
As it happenedDelivered in the final reportNow it’s in a government review.
With labelsLeft out. The check found nothing, so the claim never reaches the finished document.Finished documents carry no labels. They only carry claims that passed the check.

The judge never wrote these words. The quote is as printed in the first version of the report, as reported by the Australian Financial Review.


How the Labels Could Have Helped


What Would Have Had to Be True

The four conditions every case shares: the AI tool used the labels; it labelled its own work correctly (the weakest link: in our tests, AI sometimes mislabels its own work); the label stayed on when the text was copied; and someone owned a rule that unchecked claims don’t go further. In this case:


What Already Existed


What the Labels Wouldn’t Have Caught


Further Reading

Also listed in the AI Incident Database (#1193).

Next: All case studies · How the labels work