Build1 distinct publisher3 min readPublished
A cross-engine study of 1,487 local-service prompts found ChatGPT and Gemini rarely pick the same business. The within-engine number is the one that matters more.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Divide 1,487 queries by 50 metros and 10 verticals and you get roughly three prompts per cell [9]. That arithmetic sets a ceiling on how much of the report can be read as category insight. The finding that personal injury law, dentistry and auto repair show different citation behaviour [6] is pooled across markets by necessity, because no single metro-vertical pair has enough runs behind it to separate a pattern from one query's luck. The headline figure survives better, since it pools everything: about 62 of the 1,487 identical prompts produced the same top business in both engines [10].
The comparison worth sitting with is inside one engine, not between two. Steady Demand's benchmark puts Gemini's top-match stability at roughly 7%, against about 90% for Google's local pack [3]. Cross-engine agreement came in at 4.2% [1]. Two different assistants therefore agreed about 60% as often as Gemini agreed with its own earlier answer [13]. The gap between ChatGPT and Gemini is only a little wider than the gap between Gemini and Gemini.
The grounding figures point the same direction. Re-running queries word for word, cited sources matched about 40% of the time [4], which means roughly 60% of the evidence base moved between runs [11]. The report calls that grounding drift [4].
This complicates the study's own prescription. Reporting each engine separately [14] is necessary, but it does not produce a per-engine score either, because the per-engine answer is not stable enough to be a score. What can be measured is a rate: share of N runs in which a business is named, with N large enough to have an error bar. The study's design cannot deliver that at vertical level with three prompts per cell [9], which is a limit on the research, not on the method.
The asymmetry in sources is the part that translates into work. Gemini leaned on business websites; ChatGPT drew more on Reddit and directories [5]. Steady Demand reads that as a signal that clear on-site service information matters more in Gemini, while listing accuracy and public discussion matter more in ChatGPT [5]. That is inference from citation counts, not a tested intervention, and the mix changes by category [6], so the two work orders it implies are bets rather than fixes. National platforms including Angi, the Better Business Bureau and Reddit recur across metros [8], which at least gives a short list of things worth auditing regardless of which engine you are chasing.
One boundary is worth stating plainly, because the report states it: the study counts citations and top-name outcomes, not conversions or ranking quality [7]. Nothing here establishes that being named by an assistant produces a booking. Anyone moving spend on the strength of a 4.2% divergence is buying instrumentation for a channel whose demand has not been measured in this data.
Ranked by verification strength, evidence, and original report placement.
The study measures citations and top-name outcomes, not conversions or overall ranking quality, and does not prove AI responses drive more leads than conventional local search.
The research, published by Steady Demand in its AI Citation Ledger, examined 1,487 queries across 50 U.S. metropolitan areas and 10 service verticals, using prompts such as "best plumber near me" and tracking both the businesses named and the sources used to ground responses.
The report recommends testing engines separately, using representative prompts with a defined location, repeating the same prompts over time, recording both the business named and the cited sources, and segmenting results by category and market.
The report says a one-off screenshot of a favourable recommendation is weak evidence of durable visibility, and that the same applies to a negative or missing mention.
A cross-engine study of local-service searches found that ChatGPT and Gemini named the same top business in only 4.2% of identical queries.
In the study's benchmark, Google local-pack top results reappeared about 90% of the time, compared with roughly 7% top-match stability for Gemini's AI-generated results.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One secondary account of an unlinked vendor study
Every quantitative claim in the cluster traces to a single dev.to write-up of a study published by Steady Demand, with no link to, excerpt of, or independent replication of the underlying AI Citation Ledger data. The reported design (1,487 queries, 50 metros, 10 verticals) is internally consistent and the article states its own scope limits clearly, which is why this is not near-zero. But the run protocol, alignment definition, engine versions and per-cell allocation are all missing, and the design averages only about three queries per metro-vertical combination, so the headline percentages cannot be verified or reproduced from supplied material.
No adoption data supplied
The cluster contains a measurement exercise, not adoption evidence. It reports how two assistants answered local-service prompts but supplies no usage disclosures, traffic or lead volumes, deployment counts, or evidence that businesses have adopted the recommended cross-engine monitoring practice. The article states outright that it measures citations and top-name outcomes rather than conversions, so no adoption level can be scored without inferring facts the sources do not contain.
Precise numbers outrunning verifiable backing
Moderately overstated rather than egregiously so. The article carries genuine restraint — it says the study does not prove AI answers drive more leads, does not validate any markup or directory tactic, and warns that greater local-pack stability does not make a channel better. Against that, a two-decimal cross-engine figure and a headline framing are presented as settled findings while the underlying ledger is unlinked, per-vertical claims rest on thin per-cell sampling, and the piece closes by selling a cross-engine visibility scan. The method advice is fairly stated; the certainty implied by the percentages is what runs ahead of the evidence.
Measurement vendor plus publisher CTA
Two stacked commercial interests are visible in the source itself. The study is published by Steady Demand under a branded product name, the AI Citation Ledger, and its finding — that AI visibility is unstable and must be tracked repeatedly per engine — is precisely the conclusion that creates demand for ongoing measurement. The article then closes by promoting Scalevise's AI Visibility and GEO Checker and an AI Visibility scan for SMBs. Neither interest is disclosed as a conflict. Score is not higher because the piece still publishes limitations that cut against a maximal sales pitch.
Low-moderate: direction plausible, magnitudes unverified
Confidence is limited by single-source dependence, absent primary data, and unstated measurement definitions, and further reduced by the commercial incentives attached to the finding. What holds up reasonably well is the qualitative direction — AI assistant answers vary across engines and across repeat runs more than a conventional local pack — and the methodological prescription, which is self-consistent and testable by any reader. The specific percentages, the per-vertical and per-metro patterns, and the 13x stability ratio should be treated as unconfirmed.
build
ChatGPT tops Google's paid-click share at 4.75%, and growth teams should reprice the auction1 distinct publisher
invest
Nearly half of ChatGPT's advisor citations were advisors' own websites1 distinct publisher
build
AI referrals are 1.08% of visits, and the referral log is the wrong instrument1 distinct publisher
build
Under 30% citation overlap between engines makes pooled AI visibility scores unbuyable1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 24, 2026