Build1 distinct publisher3 min readPublished
Aggarwal et al. report visibility gains of up to roughly 40% inside generated answers, but the benchmark supplies the source documents, so the number scores selection and leaves crawl and index eligibility upstream of anything you can format.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
GEO-bench, from the 2023 paper that coined the term [12], pairs each user query with source documents [1]. That construction is the whole reading of the headline number. The document sits in the input set before the generator writes anything, so what gets scored is selection among candidates. The dev.to author says as much: it is an improvement in how likely an already-plausible page is to be selected and quoted, a shift among candidates already in contention rather than a promotion from invisible to authoritative [2].
The gate upstream of that is binary, and Google's own documentation says where it sits. AI Overviews and AI Mode draw on the same index and the same quality systems as classic Search, and there is no separate AI submission process [3]. A page has to be indexed and eligible to be shown with a snippet before formatting work can reach it at all [4]. The cases the author reports investigating were mostly that shape: a page that renders nothing without JavaScript, or a bot rule someone switched on in a CDN eighteen months ago and nobody remembers [5]. FAQ markup leaves both problems untouched.
The budget arithmetic follows from Pew's panel. Clicks on a link inside an AI summary run at 1% of visits [6]. Grant the benchmark's 40% in full and apply it there: 1% becomes 1.4%, a gain of four tenths of a percentage point [2]. The traditional-result click rate in the same panel falls by seven percentage points when a summary is present [1], which is about seventeen times the best case gain from the citation lift [3].
For those numbers to belong in one sentence, GEO-bench's visibility metric would have to map onto citation clicks in Google's summaries, on a document mix resembling yours, with the same query types. The paper's second finding, that which technique works varies significantly by domain [7], is the paper itself confirming that the transfer is uneven across domains.
The piece is headlined as a licensing decision with a formatting problem attached [8], and the argument underneath is about admission rather than contract terms [3][4][5]. Those are the same layer viewed from two ends. A robots rule and a content deal both answer whether you are a candidate; one you configure, the other you negotiate. Formatting only orders the set you were already admitted to.
So the sequence I would fund, in my context, starts with an eligibility check, then structured data, because FAQ, HowTo, Organization and Article markup give a retrieval system a fact to extract rather than infer [9]. Then narrow pages with the answer near the top of each section, because these systems lift a passage, and a page that spends three paragraphs warming up leaves the model choosing between the preamble and nobody [10]. Then a dull afternoon making the description of the business identical in every directory, because answer engines cross-reference sources rather than trusting one, and three different descriptions get attributed less confidently than three matching ones [11].
Ranked by verification strength, evidence, and original report placement.
Aggarwal et al. introduced both the term GEO and GEO-bench, a benchmark of diverse user queries paired with source documents, and demonstrated visibility improvements of up to roughly 40% inside generated answers.
The dev.to author argues the roughly 40% is an improvement in how likely an existing page is to be selected and quoted, applied to a page that was already a plausible candidate, and that nothing in the paper suggests you can format your way from invisible to authoritative.
Google's documentation states that AI Overviews and AI Mode draw on the same index and the same quality systems as classic Search, and that there is no separate AI submission process.
If a page is not indexed and eligible to be shown with a snippet, no amount of GEO work reaches it.
A Pew Research Center panel study of 68,879 searches found that when an AI summary is present, 8% of visits result in a click on a traditional result, against 15% when no summary appears, and that clicks on a link inside the summary itself are 1%.
The paper's second finding is that which technique works varies significantly by domain: the tactic that lifts a technical comparison page is not the tactic that lifts a local service page.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · September 1, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Thirty runs of one prompt put a 34-point error bar around a brand's mention rate1 distinct publisher
build
Under 30% citation overlap between engines makes pooled AI visibility scores unbuyable1 distinct publisher
product
Google's new Preferred Sources button hands publishers the recruiting job2 distinct publishers
build
Cloudflare's one-click AI block names GPTBot, not the bot that decides if ChatGPT cites you1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Named primaries, single relay
Three outside authorities hold this story up — the Aggarwal et al. benchmark, Google's own documentation, and Pew's 68,879-search panel — and a reader could go and check all three. None of them is in front of us; each arrives as a sentence in one practitioner's post. The claims the piece is most emphatic about, meanwhile, are the ones nobody measured: markup mattering more than it used to, passage-shaped writing winning, identical descriptions beating numerous ones.
Measurement adopted, practice unobserved
What has demonstrably happened is measurement, not uptake. GEO-bench exists and Pew ran a real panel at scale, so we know how searchers behave and how the techniques score in a lab. What nobody here counts is sites: how many restructured pages, how many set Google-Extended, how many changed anything and saw a citation appear. The only field evidence is a consultant's memory of client audits.
Deflates its own headline number
Most stories carrying a 40% figure are selling it; this one spends its length explaining what the figure cannot buy — selection among candidates, not discovery, and a citation that converts to a session about one per cent of the time. That earns a negative reading. It is pulled back toward zero by the one stretch: turning 40% into 0.4 points of visits requires benchmark visibility and Google summary clicks to be the same measurement, which the author flags as an assumption and then builds a seventeen-to-one conclusion on.
Practitioner warning, undisclosed book of work
The author tells you not to buy GEO on traffic projections while citing the client cases that qualify him to say so — a warning and a credential in the same breath. Nothing here says who pays for those audits, and a developer community post carries no editorial layer that would ask. The pressure runs mildly against hype rather than toward it, which is unusual enough to note.
Checkable, not yet checked
Two of the three external numbers would probably survive verification; the advice around them is not the kind of claim that can be verified at all. One publisher, no reply from the sellers the piece argues against, no view of answer engines other than Google's, and the conclusion most likely to be quoted — seventeen times — is arithmetic resting on an equivalence the author admits he has not shown.