Build1 distinct publisher3 min readUpdated
A one-afternoon test splits AI visibility in two: name recall lags funding and press by years, while category retrieval is already working for products the model cannot describe.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
A developer ran 48 AI-native products through a two-part recognition test in a single afternoon: a model with no web access was asked "What is [brand]?", then a search-grounded model was asked the category question a buyer actually types, with the brand name withheld [15][16]. The model described exactly 4 of the 48 correctly, and 28 of them appeared in the category answers, many at number one [1][3]. The gap between those two numbers is the operational point, because the two halves are fixed on completely different timescales.
The 44 misses came back with near-identical text. For Decagon, the model returned "I do not have reliable information about the software product named Decagon (decagon.ai) and cannot provide accurate details about its features, use cases, or pricing" [2][5]. Decagon is a funded customer-service AI company and Mercor is a talent marketplace with a valuation most founders would envy; neither could be described from its name [6]. The four that passed were ElevenLabs, Suno, Runway and Cursor [4]. That is 8 percent recall by name against 58 percent presence by category, roughly seven times as many appearances when the question is about function rather than brand [1][2][3].
The mechanism is unglamorous. A model's memory of names is built from training data, which is mostly the open web talking about you: Wikipedia, hundreds of G2 and Capterra reviews, a Crunchbase profile, Reddit threads, press. A product shipped last quarter has a homepage and maybe a Product Hunt launch, so there is almost nothing to recall [8]. Recall therefore tracks funding and press, both of which take years, which makes it a lagging indicator rather than a verdict on the product [7].
Three things stack on top. Training cutoff means anything published after the last run is invisible to memory until the next one [9]. Generic names actively hurt: Sierra, Wonder, Peek, Ray, Brew, Marx and Nora were all in the sample, and asked "what is Brew," the model has nothing to separate the email tool from the drink [10]. And a JavaScript-only site sends an empty body to crawlers including GPTBot and PerplexityBot, which do not execute JS, so the one source unambiguously about you says nothing [11]. The first two are slow work; the third is an afternoon, and it feeds the retrieval path that is already producing results [17].
Those results, as reported: Decagon, unknown by name, ranks first for AI customer service automation next to Ada, Intercom's Fin and Zendesk AI [12]. Lindy ranks first for AI automation platforms, ahead of Zapier and Make [13]. Sierra places fifth for customer experience platforms [14]. Half the sample was invisible by name and visible by function simultaneously, which is 24 of the 44 unrecognised products, or 55 percent of them [18][4].
Treat the numbers as one operator's single pass. The write-up does not name the models used, does not report repeat runs, and the category prompts were written by the tester rather than harvested from real buyers [19]. Search-grounded rankings are the volatile half of this, so a placement shown once is not a placement shown to be stable.
What to watch: whether fixing server-side rendering measurably changes category placement on the same prompts, whether the products in the 28 hold position across repeat runs, and whether any of the 44 crosses into name recall after the next training refresh [9][3][1].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
In a test of 48 AI-built startups, a language model with no web access described exactly 4 of them correctly.
The other 44 products came back with the same response: 'I do not have reliable information about the software product named [X].'
When asked the category question instead of the brand name, 28 of the 48 products showed up, many at number one.
Only four products passed the name test: ElevenLabs, Suno, Runway and Cursor.
The model's response for one product read: 'I do not have reliable information about the software product named Decagon (decagon.ai) and cannot provide accurate details about its features, use cases, or pricing.'
Decagon is a funded customer-service AI company and Mercor is a talent marketplace that has raised at a valuation most founders would trade a kidney for; the model could not describe either from its name.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Concrete but unreplicated single-author test
The source reports specific counts (4 of 48 named, 28 of 48 surfaced by category), a verbatim model response, and named placements against named incumbents, which is more than assertion. But everything rests on one self-published run: the two systems are described only as 'a model with no web access' and 'a search-grounded model' with no vendor, version or configuration, there are no repeat runs or prompt variations, category prompts were written by the tester, and no independent source in the cluster corroborates any figure. The supporting technical claim about GPTBot and PerplexityBot not executing JavaScript is stated without logs or documentation.
No adoption or usage data supplied
The cluster contains one informal recognition test and no deployment, usage, pricing or traffic disclosures. Category ranking positions describe what a model answered, not that any buyer acted on it, and the source explicitly does not link placement to traffic or revenue. There is no basis to score adoption without inferring facts the material does not contain.
Findings generalized beyond a one-run test
The piece is unusually candid about its limits - it discloses the one-afternoon run and declines to name failed products - but the framing still moves from a single unnamed-model pass to general rules: that recall lags funding and press by years, that category share of voice is 'the metric tied to revenue', and that a same-afternoon page fix decides category visibility. None of those causal claims were tested in the described run, and no revenue or traffic link is shown, so the interpretive reach exceeds the evidence by a moderate margin rather than a large one.
Author is tester, interpreter and prescriber; no disclosure
Observable from the text alone: the same author designed the prompts, ran the test, chose which results to name, and then prescribes the fix, publishing on a personal developer blog with a search-style headline ('Why Doesn't ChatGPT Know Your Startup?') and a coined metric ('share of voice in AI answers'). That structure rewards a striking result. The supplied material contains no sponsorship, vendor affiliation or commercial disclosure either way, so this is scored on the self-published-methodology pattern only, not on any established financial interest.
Directionally plausible, quantitatively soft
Confidence is limited by structure rather than by contradiction: nothing in the cluster disputes the account, but nothing corroborates it either, and the single source withholds the model identities and repeat runs that would make the counts checkable. The qualitative direction - that young products are weakly represented in model memory while retrieval can still surface them - is coherent with the stated training-data and cutoff mechanics, so it deserves moderate credence; the specific percentages and the causal prescriptions do not.
build
Three files, three contracts: robots.txt, sitemap.xml and llms.txt are not rivals1 distinct publisher
build
AI-written code fails the same four ways, and every gate you own reports green1 distinct publisher
build
Notion's agent stack is live, not slideware, and it only changes one of your decisions1 distinct publisher
build
A 5x publishing increase cost one site 1,000 indexed pages and every impression1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 17, 2026