Science1 publisher3 min readPublished
Only 41.9% of Google AI Overviews had every claim backed by the pages they cited
WashU researchers checked 98,020 claims in Google AI Overviews and found only 41.9% of overviews fully backed by the pages they cited. Most single claims held up, but at about 13 claims per answer, the citations cannot vouch for a whole overview.
The Scientist · Science desk
What happened
- The team ran 55,393 Google searches on trending topics between March 13 and April 21, 2026, saving each overview, its citations and the cited pages.
- Of the domains the overviews cited, 29.8% never appeared in the first-page search results, which the researchers take as a sign of a separate source-selection process.
- Question-form searches produced an overview nearly 65% of the time, compared with 9.5% for other queries.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- exposure A reader who treats an overview's citation list as proof the answer checks out could be wrong in some detail in up to 58.1% of answers in this sample.
- decision Anyone scoring AI answer quality has to pick a unit: the failure ceiling is 11% per claim but 58.1% per answer, and quoting only the first understates what a user meets.
- constraint Checking an overview against Google's ordinary ranked results is an incomplete test when nearly 30% of the domains it cites are absent from page one.
- exposure People who type plain questions meet overviews nearly seven times as often as keyword searchers, so the grounding gap falls hardest on them.
Which denominator you choose decides how the overviews look. Counted claim by claim, they mostly hold up: about 89% of the 98,020 verifiable claims were clearly or broadly supported by the pages Google cited [9]. Counted answer by answer, they do worse. Only 41.9% of overviews had every verifiable claim supported by the available cited text [12], so 58.1% contained at least one that was not [1].
The distance between those figures comes from how much each overview says. Overviews appeared on 13.7% of the team's 55,393 searches [2][3], or roughly 7,600 answers [2]. Dividing the 61,212 citations collected [6] by that count gives about eight per overview, matching the average the authors report [8][3]. Dividing the claims the same way gives about 13 checkable statements per overview [4].
The paper, by Jacob Montgomery, Umar Iqbal and Haofei Xu of Washington University in St. Louis, is due at the ACM Internet Measurement Conference in October 2026 and is posted on arXiv [1]. The authors call the 11% failure rate an upper limit [11]. Their pipeline could not fully capture Reddit threads, forum posts or YouTube videos, and fast-changing pages such as weather forecasts and school closings may have been updated between the overview and the capture [11]. Of that 11%, 7% were claims the team could not find in the cited text it collected and 4% contradicted the cited source [10]. A missed video transcript would most likely show up in the first group, and a claim absent from a scraped page is not necessarily false. The contradictions are the firmer figure, about 3,900 claims [5], though a stale forecast could produce some of those too.
The source-quality comparison has a sensible control. The team set the cited domains against the first-page results shown for the same searches. The cited domains scored higher on average credibility and leaned less on user-generated content [6]. The two sets also diverged: 29.8% of cited domains did not appear on the first page at all [7]. The researchers say that points to a source-selection process distinct from Google's conventional ranking [7].
Who meets an overview depends heavily on phrasing. Nearly 65% of question-form queries produced an overview, against 9.5% of other searches [4], a ratio of nearly seven to one [6]. Overviews appeared for 46.1% of hobby and leisure searches and 39.9% of science queries, but only 7.5% of political ones [5]. The WashU account says that spread suggests Google's triggering logic relies on undisclosed editorial discretion [5]. "Google is now writing answers, not just ranking them, and it does this for billions of searches," Montgomery said [13].
The thing this doesn't tell you is whether an unsupported claim is wrong about the world. The audit tests whether a sentence matches its own citations, which is a narrower test than accuracy. The queries came from trending topics searched between March 13 and April 21 [2]. The published account does not break the grounding figures down by topic, so a cholesterol answer and a boiled-egg answer could fare differently. In my view the evidence supports a specific form of the worry. An overview's citations show where it drew from, and in as many as 58.1% of answers at least one sentence went further than the available cited text [1].
What to watch
- Whether the full paper breaks grounding rates down by topic, especially the health and science queries that trigger overviews often.
- A re-run that captures Reddit, forum and YouTube content, which would show how much of the 7% 'not found' group is a measurement gap.
- Any change in how Google selects or displays overview citations after the October 2026 ACM Internet Measurement Conference presentation.