Build1 distinct publisher3 min readPublished
A Search Engine Land analysis published in February 2026 found that repeating the same ChatGPT prompt returns different brands, which makes visibility a proportion you sample. The arithmetic of that proportion sets how many runs your report needs.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Start with the arithmetic, because it sets the budget. Run one prompt 30 times, count 10 answers that mention your brand, and the point estimate is 33.3%. Put a 95% Wald interval on it, p plus or minus 1.96 times the square root of p(1-p)/n, and the range is 16.5% to 50.2%, a width of 33.7 points [2]. Keep the same hit rate and take 300 runs instead, and it tightens to 28.0% to 38.7%, a width of 10.7 points [3]. Width scales as one over the square root of n, so halving it costs four times the calls [4].
Then multiply by surfaces. The writeup names five distinct answer surfaces [1], and it says the research ran prompt sets dozens to hundreds of times [3]. At 100 runs each, that is 500 observations for one prompt in one reporting period [5], before you have measured a single competitor. A single screenshot documents one run, not a measured rate with confidence bounds attached.
There's an assumption underneath the interval worth checking. The dev.to summary reports that the same prompts returned changing brand recommendations [3] but does not say what generates the variance, and that gap matters, because a binomial interval assumes the runs are independent draws from one fixed distribution [6]. If a model version rolls out across your sampling window, or a retrieval index refreshes mid-week, consecutive runs are correlated and the interval you print is narrower than the uncertainty you actually have. The cheap defence is already in the source's governance list: record collection dates [9], then plot the daily rate rather than the pooled one, so a step change shows up as a step instead of dissolving into the mean.
Per-surface separation is not a presentation preference either. The source's position is that each surface is its own measurement environment, and that blending ChatGPT, the Google experiences, Perplexity and Gemini into one unexplained number hides the differences you were paying to see [10]. Same logic applies to the counting rule. "Mentioned" and "recommended" are different Bernoulli variables [9], and if you switch between them between quarters, your trend line is measuring your own definition.
This is where someone else's number stops transferring. A published mention rate is a claim about a specific prompt list, a specific surface, a specific sampling window, and a specific rule for what counted as a hit [9]. Changing the prompt wording changes the population, and changing the surface changes the instrument [10]. The source is explicit that results from one surface should not be assumed to represent another [7], which is a polite way of saying that a vendor's cross-platform score is not comparable to your own unless you can read their method.
What follows is narrow and practical. A single run is fine evidence of an opportunity or a content gap [6], and useless as a position. If two brands' intervals overlap, the gap between them is not a budget input [8]. And a period-over-period change is unfalsifiable unless the prompt file, the run count, the dates and the counting rule shipped with the percentage [7][9]. Without the method sitting next to the number, the number is just a screenshot with more decimal places.
Ranked by verification strength, evidence, and original report placement.
A brand appearing in a ChatGPT answer is not the same as holding a stable search ranking; repeated runs of the same prompt can surface different brands and sources, making AI visibility a variable that must be sampled rather than captured in a one-off report.
Running the same prompts dozens or hundreds of times produced changing brand recommendations.
The report places ChatGPT alongside Google AI Mode and AI Overviews, Perplexity, and Gemini as surfaces where brand visibility should be treated as a moving target.
The methodological guidance highlighted in the research is to use repeated runs, compare across surfaces, and report results with confidence intervals.
A single screenshot can show that a brand appeared but cannot reliably show how consistently it appears; a one-time result is useful evidence of an opportunity or content gap, not a definitive measure of standing.
The reported variability undermines four common conclusions: a single mention does not establish dependable visibility; an apparent competitor lead may reflect a sample of one; results from one AI surface should not be assumed to represent another; and changes between reporting periods may reflect sampling variation unless the testing method is consistent.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 31, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Under 30% citation overlap between engines makes pooled AI visibility scores unbuyable1 distinct publisher
build
Reddit's ChatGPT citations fell 86% in four days, and the change was OpenAI's, not Reddit's1 distinct publisher
build
Two assistants agreed on the top plumber 4.2% of the time. One agreed with itself 7% of the time.1 distinct publisher
build
Cloudflare's one-click AI block names GPTBot, not the bot that decides if ChatGPT cites you1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One retelling, no underlying data
Everything this story asserts about how ChatGPT actually behaves comes from a single dev.to post summarising a Search Engine Land analysis we never see — no link, no prompt list, no run counts, no distributions. 'Dozens or hundreds of times' is the closest thing to a method. What does hold is the part that needs no source: the interval arithmetic around a sampled mention rate, which is ours and is checkable line by line.
Not established
Nobody says how many teams have actually moved to repeated sampling, what it costs them, or whether any reporting standard has shifted. The one datum resembling uptake is Scalevise advertising its own checker at the end of the post, and a seller's intent is not a user's behaviour.
Modest tilt, mostly at the base
The recommendation is deliberately deflationary — sample more, publish error bars, stop reading one screenshot as a ranking — so there is little for a headline to inflate. The overstatement sits underneath: the research is described as documenting the behaviour, in a piece that never shows it, and the post's own warning about dashboards implying unearned precision applies to its own sourcing. The sales close nudges it a little further.
The advice and the product match
The piece argues that snapshots are worthless and that credible answers require repeated runs, per-surface splits and documented governance — then names Scalevise's AI Visibility and GEO Checker as the way to get them. That is not disqualifying; the statistics are true regardless of who says them. It does mean the recommended volume of work is also the billable surface, and no disclosure of that alignment appears anywhere in the text.
Low on the facts, high on the maths
Split the story in two and the confidence splits with it. That AI answers vary between runs is plausible and consistent with how these systems work, but here it is one publisher's summary of an analysis nobody in our coverage has opened. That thirty runs cannot distinguish a 20% brand from a 45% one is not a matter of trust at all. We would raise this materially on sight of the original data, or on any second party reproducing the test.