What was in the report

In July 2025 Deloitte delivered an assurance review of the Targeted Compliance Framework to Australia's Department of Employment and Workplace Relations. The framework is the system that decides when welfare recipients get penalised for missing obligations, so the review was meant to check whether the department's automated decisions were lawful and well founded. The contract was worth about A$440,000.

The problems were found by Chris Rudge, a researcher at the University of Sydney who works on welfare law. Reading the report, he noticed citations to academic work he could not find and to academics who did not appear to have written the papers attributed to them. Then he found a quotation attributed to a Federal Court judgment in the Amato case, one of the robodebt decisions, which did not appear in the judgment. A quote invented and put in a judge's mouth, in a report about whether the government's own automated decisions were lawful, is about as pointed an irony as this field produces.

What the fix looked like

Deloitte replaced the report with a corrected version. The revised document removed the references that could not be verified, corrected the court citation, and added a disclosure that had not been in the original. That disclosure stated that a generative model, Azure OpenAI GPT-4o, had been used in preparing the work. Deloitte agreed to repay the final instalment of the contract. The full refund that some senators asked for did not happen.

The department's position, as quoted in the coverage, was that the updates in no way affected the substantive content, findings and recommendations of the report. We understand why a department that has to act on those recommendations would say that. But it is a claim that deserves scrutiny rather than acceptance. If the references that supported an argument were invented, the argument was not supported. It might still be right, but the reader has lost the ability to check, which is what references are for.

How this happens inside a firm

None of the reporting suggests a single author sat down and typed a model's output into a final deliverable without looking. The more likely story, and the one we have seen in other organisations, is a chain. Someone uses a model to draft a literature section or to tidy a paragraph. The model produces text with the shape of a citation, with a plausible author list and a plausible journal. A reviewer checks the prose reads well and the argument makes sense. Nobody's job is to open each reference and confirm it exists, because until recently references written by a human almost always did.

That is the structural point. Citation verification used to be cheap because the error rate was near zero. Models moved the error rate to something noticeable, and a control that was never formalised turned out not to exist. The report got through a professional services quality process, and through a government department's acceptance process, with an invented quotation from a court in it. Both sides had checks. Neither check was designed for this failure.

Why this became a contract term

The refund settles one contract. The consequence we expect to matter most is the disclosure. Once a government client has been embarrassed by a model's output in a paid deliverable, its procurement teams start asking two questions in every future engagement. Was a model used, and how was the output checked. The first question already appears as a disclosure line in the corrected Deloitte report. The second is harder to answer honestly, which is the point of asking it.

A serious answer looks like a verification step that is written into the process, with a named owner. Every citation is resolved to a real document before the draft leaves the team. Every quotation is checked against the primary source, not against a model's memory of it. Any text drafted with a model is flagged in the working file so that reviewers know where to look harder. None of this is technically difficult. The work is administrative, and firms have resisted it because it is slow and because it exposes how much of the drafting is now done by the tool.

What we would want to see next

We would like the department to publish both versions of the report side by side. A diff between the original and the corrected document is the single most useful artefact this episode could produce, because it would show exactly what kind of content the model invented and where in the argument it sat. That would tell other buyers what to look for, and it would let researchers measure the thing rather than argue about it.

We would also like to see contracts start pricing the verification work explicitly. If a A$440,000 engagement now needs a person to resolve every reference, that person's time should be in the quote, and a firm that leaves it out should be assumed not to be doing it. The refund settled one contract. The habit of asking for evidence that citations were checked is the change that lasts.

Sources

  1. The Guardian, Deloitte to pay money back to Albanese government after using AI in $440,000 report
  2. Wikipedia, Hallucination (artificial intelligence), section on the Deloitte report
  3. Hacker News discussion of the Guardian report