Leadership1 distinct publisher3 min readPublished
Guardian Australia checked the references in every submission made to the current parliament against academic databases, and some of the documents that failed the check had already been quoted in committee reports.
The Board Room · Leadership desk

Compiled by The Board RoomSomething wrong?How this is made
The asymmetry here is between the cost of producing a plausible citation and the cost of checking one. The useful part of this story for anyone who runs an intake process is the method: extract every reference from every document, match them in bulk against academic databases, then put a person on the documents where the match rate collapses [1]. That is a screening design rather than a research project, and it is available to any organisation that accepts written submissions.
It also produces a floor, not a count. The Guardian calls its own figure conservative, because reference matching only catches documents that cite anything and misses text a model wrote without citations [11]. A second marker in the same corpus indicates how much wider the usage is: more than 100 papers carried ChatGPT url tags that the platform inserts automatically [4], roughly 2.6 times the number of submissions the citation check flagged [16]. The gap between those two numbers is the part a citation audit cannot see.
What changed recently is the disappearance of the free check. Looking up an invented reference used to return nothing, and the nothing was the answer; Google's AI summary will now sometimes describe a fake paper as though it were genuine [9], and both that summary and ChatGPT will in some cases cite the submission containing the fake reference as their source [10]. A reviewer who searches and finds corroboration has been handed the appearance of verification.
The skeptic's version of this comes from the submitter. Drilldown Reports told Guardian Australia that it used AI in its research process, identified the errors itself, and filed a follow-up submission after the deadline for corrections had passed, and its spokesperson said "this was a human error, not AI" [7][8]. Take that at face value and the exposure is unchanged, because the misattributed reference stays in the file and the window for fixing it was set by the process calendar rather than by the date the error surfaced. The researcher's objection, in Divna Haslam's account, is that this devalues rigorous work and that misleading material can travel into policy, particularly in the domestic violence area [15].
The board-deck version is that this is a process problem in one parliament. It is incomplete because the shape is generic: unpaid written input, a hard deadline, a reviewer with no verification budget, and an output that quotes the input. RFP responses, public consultations and grant applications all sit in that shape. The committee chair's stated allocation of cost is to accept evidence on face value and then interrogate it through the inquiry [13], which is defensible when submissions are expensive to write and weaker when they are cheap to generate. Christian Downie of the Australian National University describes the end state as parliamentarians deciding on evidence that does not exist [12].
The trade-off runs in both directions. Screening citations at intake costs staff time and will flag honest submitters with sloppy referencing habits, while skipping it means your own report can quote work nobody ever published. This quarter the decision is narrower than the problem, and it concerns which documents you intend to quote, because those are the citations you own once they appear over your name.
Ranked by verification strength, evidence, and original report placement.
Guardian Australia built a computer program that extracted references from all inquiry submissions made to the current Australian parliament and checked those references against several online academic databases; documents flagged as having a high proportion of references not matching anything were then manually checked.
The Guardian described its findings as a conservative estimate of the amount of AI-generated material in political processes, because the method only identifies submissions with incorrect citations and would not find documents that used AI to generate text without any references.
Guardian Australia found at least 39 submissions to Australian political inquiries containing what appear to be hallucinated references.
In some cases, Australian parliamentary committee reports have cited submissions in which the majority of sources appear to be AI-generated hallucinations.
More than 100 papers submitted included ChatGPT url tags in reference links, which are automatically added by the platform in its results.
One document submitted to an inquiry into family violence and suicide included a hallucinated reference attributed to Divna Haslam, a University of Queensland associate professor and clinical psychologist, and misstated her team's research findings.
Follow any of these and your For You feed starts watching them — no settings page required.
leadership
A $3.48m evidence base, six bad citations, and a denial that did not survive the metadata1 distinct publisher
science
Claude's watermark is a compliance artefact, not a cheating detector1 distinct publisher
science
Text watermarks land on 2 December. The detection they imply does not.1 distinct publisher
invest
A Connecticut judge just priced prompt injection: no fine, no e-filing2 distinct publishers
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One newsroom's audit, corroborated case by case
The findings and the measuring instrument have the same author: Guardian Australia wrote the reference-extraction program, set the threshold for what counted as suspicious, and did the manual review, and none of that code or the flagged list is published for anyone to rerun. What lifts it well above assertion is that the individual cases hold up under their own names — Divna Haslam confirming a citation in her name she never wrote, Nicole Gurran the same, and Drilldown Reports admitting on the record that AI was in its research process. The counts remain unreplicated; the specimens do not.
Counted artifacts, missing denominator
These are not survey estimates. Thirty-nine submissions failed database matching and more than a hundred documents still carried ChatGPT's own link tags, which is uptake you can point at. Two things hold the number down: Guardian Australia never says how many submissions the current parliament received, so 'flooded' cannot be sized as a rate, and the reach stops at Australian federal inquiries. The detail that makes this consequential rather than merely present is that committee reports went on to quote the flagged documents.
Framing runs ahead of the arithmetic, slightly
'Flooded' is doing more work than the numbers can bear while the total submission count stays unstated, and the fabrications are consistently described as what 'appear' to be hallucinations — the step from 'this reference matches nothing' to 'a model invented it' is inference, and Drilldown Reports contests exactly that step. Pushing the other way, the paper volunteers that its net only catches cited fakery and misses AI-written prose without footnotes, and the ChatGPT-tag count sits well above the flagged total. Overstatement in the verb, not in the findings.
Everyone quoted has a stake
Trace who gains from each sentence. Guardian Australia holds the exclusive and says so, which rewards the widest available framing of its own audit. Drilldown Reports, the only named author, has every reason to file this under 'human error, not AI' — its spokesperson says as much while calling the story the 'AI fear machine'. The committee chair's reply defends a screening practice he presides over. The academics are protecting the value of citation, which is their currency. And the two companies whose products both generated the fake references and later summarised them as real were not asked to comment at all.
Act on the pattern, not on the number
Specific enough to be useful: named researchers who confirmed fabrications carrying their names, an author who admitted AI use, a documented method, and a mechanical fingerprint in the ChatGPT link tags. Thin where it counts: one publisher, no replication, unpublished detection code, an unknown base rate, and a causal attribution the accused party disputes. Institutions should treat the screening gap as established and the 39 as a floor of one newsroom's making rather than a census.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 31, 2026