Leadership1 publisher3 min readPublished
A 3,000-run test reproduced the same AI brand list in fewer than 1 percent of repeats
Vendors sell AI visibility as measurement, but the prompt panels behind the scores are modeled and the answers move run to run. Building the panel from your own sales calls buys provenance and narrows the sample to buyers who already found you.
The Board Room · Leadership desk

What happened
- The Interactive Advertising Bureau's August 2026 AI visibility guidance says more than 20 companies use different methodologies that can produce different answers about the same brand.
- Research across 693,509 repeat answers found that two responses to the same ChatGPT prompt shared 21.2 percent of their cited domains.
- No major AI discovery platform exposes a complete query stream comparable with traditional search-query reporting, though Otterly and Ahrefs publish how they assemble their question sets.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
- constraint A score built from one run cannot carry a quarterly target, so any team that reports one is committing to repeated observation on a panel that stays locked for the whole quarter.
- decision Leadership has to pick which flaw it prefers: a vendor panel with wide reach and undisclosed provenance, or an in-house panel whose provenance is known and whose sample stops at buyers who already made contact.
- exposure Whoever signs off on the number inherits the vendor's method, and a competitor citing a different vendor can produce a different answer about the same brand.
- capability Owning the question set changes what the subscription is for: the platform becomes an instrument a buyer can audit against its own baseline.
The prompt list is the sample frame. In a modeled panel the vendor chooses it, so the score a company reports describes a population the vendor defined. According to the Entrepreneur account, no major AI discovery platform exposes a complete query stream comparable with traditional search-query reporting; the questions in the report are generated, not recorded [1][2]. Two vendors publish how theirs are built: Otterly documents its use of Search Console data, keyword research and generated brainstorming, and Ahrefs publishes a methodology for expanding related questions [3].
Underneath the panel question sits the instability of the answers themselves. Research covering 693,509 repeat answers found that two responses to the same ChatGPT prompt shared 21.2 percent of their cited domains [7]. Put the other way, 78.8 percent of the domains cited in one answer were absent from the other [8]. In a 2026 crowdsourced study, 600 volunteers ran the same brand-recommendation prompts through major AI systems nearly 3,000 times, about five runs each, and the same list of brands came back in fewer than one run in a hundred [6][9].
Noise does average out with enough observations, and that is what the Interactive Advertising Bureau's guidance is built around: it separates directional data from decision-grade data and treats a program of fewer than 50 queries as exploratory [5]. Averaging also requires a question set that stays fixed between runs, and the guidance's distinction only holds if the panel does not drift [18]. The number on screen in a sales meeting is a single run [18].
The choice between buying the panel and building it is coverage against provenance. Sales calls, support tickets, win-and-loss debriefs and community threads carry real buyer questions in the buyer's own language, and a competitor has no access to them [10]. They also cover only the buyers who reached you, not everyone researching the category [11]. The piece's remedy is to supplement first-party questions with public category questions and keep the panel locked long enough to compare results over time [11]. It also argues the panel should span the decisions a buyer is making: discovery, comparison, risk, proof and commercial questions [12].
The pitch, in the author's telling, arrives in a fixed order: your buyers ask these questions, you appear here in the response, your competitor appears above you [17]. "The first time I asked where the question set came from, the room got noticeably less specific," the author wrote [13]. The same piece puts repeatability, source patterns and disclosed methodology above the cleanest-looking score in a demo [14].
This is one contributor's opinion column, two studies and a trade-body guidance document [19]. It does not show that a first-party panel predicts revenue, or that teams running one decide better than teams reading a vendor score. It does support a governance line the piece states directly: modeled prompt panels have their uses, and should be priced, governed and reported as modeled demand [16]. If a vendor score enters this quarter's board pack as a KPI, someone owns the explanation when it moves next quarter, and with more than 20 companies producing different answers for the same brand, that explanation begins with a methodology the company did not write [4].
What to watch
- Whether any AI discovery platform starts exposing a complete query stream. That would move panels from modeled to recorded.
- Whether IAB's 50-query floor and its directional versus decision-grade split turn up in procurement language and vendor contracts.
- Whether repeat-run stability improves as model providers change retrieval; the 21.2 percent citation overlap is the baseline to measure against.