Skip to content

Science1 publisher3 min readPublished

Two governments got AI slop in writing, and the review chain caught none of it

South Africa pulled a draft national AI policy over invented citations, and Deloitte refunded part of a A$440,000 report for the same defect. The failure was institutional, not personal.

The Scientist · Science desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • South Africa's Draft National AI Policy had to be withdrawn earlier this year after reviewers found that several citations pointed to journal articles and authors that did not exist; the mistakes were introduced by AI.
  • Deloitte provided a partial refund to the Australian government after admitting that generative AI had been used to help produce a commissioned document containing fabricated citations and quotes attributed to sources that did not exist.
  • The Deloitte report commissioned by the Australian government cost Australian taxpayers A$440,000.
  • Work produced by AI often looks polished on the surface but carries significant underlying flaws, a byproduct of a modern workplace that conditions people to prioritise the final deliverable over the actual process.
  • When the process is skipped, people lose the ability to fully grasp what the output means, making it harder to critically evaluate outcomes and to discern facts from partial truths and lies.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

South Africa withdrew its Draft National AI Policy earlier this year after reviewers found several citations pointing to journal articles and authors that do not exist, errors introduced by AI [1]. Deloitte separately gave the Australian government a partial refund after admitting generative AI had helped produce a commissioned document containing fabricated citations and quotes attributed to sources that did not exist; the report cost Australian taxpayers A$440,000 [2][3].

Both cases get filed as individual carelessness. That reading lets the organisations off. A national policy draft and a paid consultancy deliverable each move through a chain of hands, and in both cases the fabricated material survived the whole chain and only surfaced at the far end, either in external review or after delivery and payment [1][2].

The phys.org analysis argues the mechanism is a workplace that rewards the final deliverable over the process that produced it [4]. Skip the process and you lose the ability to grasp what the output actually means, which then degrades your capacity to tell facts from partial truths [5]. That is a quality-control problem with a procurement shape: the buyer inspects polish, and polish is exactly what a language model supplies cheapest.

The adoption data suggests the exposure is broad. A 2025 study by the University of Melbourne and KPMG surveyed more than 48,000 respondents across 47 countries and found two in three people use AI regularly, with more than half believing their performance benefits from it [6][7]. The primary driver of adoption was fear of missing out, at 48%, which the study links to users prioritising speed of delivery over quality of output [8]. Meanwhile 61% reported receiving no AI training at all [9], and 60% reported inappropriate, complacent or nontransparent use of AI in their workplace [10]. Two-thirds of people are using the tools regularly while under two-fifths have had any instruction in them [11].

The capability side does not close the gap. A July 2026 study by the Centre of AI Safety, a San Francisco nonprofit, found that even top-performing AI agents failed to complete roughly 85% of projects to a standard acceptable for commissioned work [12]. Put the other way, the best agents cleared the commissioned-work bar on about 15% of projects [13]. Anyone treating agent output as a finished deliverable is accepting roughly five failures for every acceptable one, before human review.

The analysis does not recommend banning the tools [14]. Its remedies are managerial: set deadlines long enough that people can actually think rather than merely go faster [15], disconnect from AI during brainstorming and structuring so the initial synthesis is human, including plain whiteboard sessions [16], and build a culture that is constructively critical of every output, human or machine [17].

Worth watching is whether the Deloitte remedy generalises. A partial refund converts fabrication from a reputational event into a priced contract defect, which is the only version procurement departments reliably act on [2]. Watch also for verification duties written explicitly into consulting and policy contracts, and for the 61% untrained figure in the next wave of the Melbourne and KPMG survey [9]. If regular use keeps climbing while training stays flat, the South African withdrawal is a template rather than an outlier.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories