Science1 distinct publisher3 min readUpdated
South Africa pulled a draft national AI policy over invented citations, and Deloitte refunded part of a A$440,000 report for the same defect. The failure was institutional, not personal.
The Scientist · Science desk
Compiled by The ScientistSomething wrong?How this is made
South Africa withdrew its Draft National AI Policy earlier this year after reviewers found several citations pointing to journal articles and authors that do not exist, errors introduced by AI [1]. Deloitte separately gave the Australian government a partial refund after admitting generative AI had helped produce a commissioned document containing fabricated citations and quotes attributed to sources that did not exist; the report cost Australian taxpayers A$440,000 [2][3].
Both cases get filed as individual carelessness. That reading lets the organisations off. A national policy draft and a paid consultancy deliverable each move through a chain of hands, and in both cases the fabricated material survived the whole chain and only surfaced at the far end, either in external review or after delivery and payment [1][2].
The phys.org analysis argues the mechanism is a workplace that rewards the final deliverable over the process that produced it [4]. Skip the process and you lose the ability to grasp what the output actually means, which then degrades your capacity to tell facts from partial truths [5]. That is a quality-control problem with a procurement shape: the buyer inspects polish, and polish is exactly what a language model supplies cheapest.
The adoption data suggests the exposure is broad. A 2025 study by the University of Melbourne and KPMG surveyed more than 48,000 respondents across 47 countries and found two in three people use AI regularly, with more than half believing their performance benefits from it [6][7]. The primary driver of adoption was fear of missing out, at 48%, which the study links to users prioritising speed of delivery over quality of output [8]. Meanwhile 61% reported receiving no AI training at all [9], and 60% reported inappropriate, complacent or nontransparent use of AI in their workplace [10]. Two-thirds of people are using the tools regularly while under two-fifths have had any instruction in them [11].
The capability side does not close the gap. A July 2026 study by the Centre of AI Safety, a San Francisco nonprofit, found that even top-performing AI agents failed to complete roughly 85% of projects to a standard acceptable for commissioned work [12]. Put the other way, the best agents cleared the commissioned-work bar on about 15% of projects [13]. Anyone treating agent output as a finished deliverable is accepting roughly five failures for every acceptable one, before human review.
The analysis does not recommend banning the tools [14]. Its remedies are managerial: set deadlines long enough that people can actually think rather than merely go faster [15], disconnect from AI during brainstorming and structuring so the initial synthesis is human, including plain whiteboard sessions [16], and build a culture that is constructively critical of every output, human or machine [17].
Worth watching is whether the Deloitte remedy generalises. A partial refund converts fabrication from a reputational event into a priced contract defect, which is the only version procurement departments reliably act on [2]. Watch also for verification duties written explicitly into consulting and policy contracts, and for the 61% untrained figure in the next wave of the Melbourne and KPMG survey [9]. If regular use keeps climbing while training stays flat, the South African withdrawal is a template rather than an outlier.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
South Africa's Draft National AI Policy had to be withdrawn earlier this year after reviewers found that several citations pointed to journal articles and authors that did not exist; the mistakes were introduced by AI.
Deloitte provided a partial refund to the Australian government after admitting that generative AI had been used to help produce a commissioned document containing fabricated citations and quotes attributed to sources that did not exist.
The Deloitte report commissioned by the Australian government cost Australian taxpayers A$440,000.
A 2025 study by the University of Melbourne and accounting firm KPMG surveyed more than 48,000 respondents across 47 countries.
The Melbourne and KPMG study found that two in three people (66%) use AI regularly, and more than half believe their performance benefits from its use.
The study found the primary driver of AI adoption was fear of missing out (48%), leading users to prioritise speed of delivery over the quality of the final output.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single secondary account, no primary documents
Everything rests on one republished opinion analysis. The two institutional incidents are stated flatly with no links to the withdrawn policy, the Deloitte report, the refund terms or any official statement, and the article's own date framing is loose ('earlier this year', 'last year'). The survey figures are internally consistent and specific enough to be checkable, but the agent-capability study is a one-line unlinked reference with an organisation name that is given only as 'Centre of AI Safety' and no methodology. The causal core of the piece — that offloading work degrades cognitive capacity — is asserted rather than evidenced.
Widespread untrained use, two documented institutional failures
Adoption of the behaviour in question is well attested even if the harm mechanism is not: a 47-country survey puts regular AI use at 66% with 61% untrained and 60% reporting misuse in their own workplace, and two named institutions — a national government drafting a policy and a Big Four consultancy billing A$440,000 — shipped AI-fabricated citations through review. That is real deployment inside consequential workflows, not pilots. The score is held below high because all of it reaches us through one secondary account and the survey measures self-reported use rather than verified organisational deployment.
Deflationary framing, but its own causal claim outruns its evidence
The article is critical of AI rather than promotional, so it carries no vendor hype. The overstatement runs the other way: the headline premise that offloading tasks costs us our brains is supported only by a perception survey and an unlinked agent study, neither of which measures cognitive capacity or output quality. Terms like 'work slop' and 'blind reliance' generalise from two high-profile incidents to a systemic cognitive decline. The two incidents themselves and the survey numbers are stated proportionately, which keeps the gap modest rather than large.
Consultancy-sourced data and an advocacy prescription
The demand-side statistics come from a survey co-produced by KPMG, a firm that sells AI advisory and governance services into the exact market the numbers describe, and the article does not disclose that interest. The cautionary example is Deloitte, a direct competitor of KPMG. The author writes in an academic-commentary format that rewards a clear prescriptive stance, and the closing recommendations advocate practices — critique groups, whiteboard sessions, longer timelines — rather than test them. None of this indicates bad faith, and phys.org is republishing rather than originating, so the score sits mid-range.
Incidents credible, mechanism and benchmark thinly grounded
Confidence is split. The two institutional failures and the survey percentages are specific, mutually consistent and of a kind that is widely reported elsewhere, so they are probably accurate as stated. The July 2026 agent study and the cognitive-degradation thesis are much weaker: one unlinked reference and one unevidenced argument, from a single publisher with no corroborating source in the cluster and no response from any named party. That combination supports moderate confidence in the factual spine and low confidence in the interpretation built on it.
science
A handful of Texan pumas saved the Florida panther. Genetics is still off the plan.1 distinct publisher
science
Most kingfish carrying the mushy-flesh parasite cook up firm, so the cook test misses them1 distinct publisher
science
H5N1 is in Australian resident birds, and overseas the first wave does most of the killing1 distinct publisher
science
A PNAS study says X's feed learns from your arguments, not your likes2 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 19, 2026