Skip to content

Build1 publisher3 min readPublished

48 startups, 4 known by name, 28 recommended by category

A one-afternoon test splits AI visibility in two: name recall lags funding and press by years, while category retrieval is already working for products the model cannot describe.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying 48 startups, 4 known by name, 28 recommended by category
Generated illustration

What happened

  • In a test of 48 AI-built startups, a language model with no web access described exactly 4 of them correctly.
  • The other 44 products came back with the same response: 'I do not have reliable information about the software product named [X].'
  • When asked the category question instead of the brand name, 28 of the 48 products showed up, many at number one.
  • Only four products passed the name test: ElevenLabs, Suno, Runway and Cursor.
  • The model's response for one product read: 'I do not have reliable information about the software product named Decagon (decagon.ai) and cannot provide accurate details about its features, use cases, or pricing.'

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A developer ran 48 AI-native products through a two-part recognition test in a single afternoon: a model with no web access was asked "What is [brand]?", then a search-grounded model was asked the category question a buyer actually types, with the brand name withheld [15][16]. The model described exactly 4 of the 48 correctly, and 28 of them appeared in the category answers, many at number one [1][3]. The gap between those two numbers is the operational point, because the two halves are fixed on completely different timescales.

The 44 misses came back with near-identical text. For Decagon, the model returned "I do not have reliable information about the software product named Decagon (decagon.ai) and cannot provide accurate details about its features, use cases, or pricing" [2][5]. Decagon is a funded customer-service AI company and Mercor is a talent marketplace with a valuation most founders would envy; neither could be described from its name [6]. The four that passed were ElevenLabs, Suno, Runway and Cursor [4]. That is 8 percent recall by name against 58 percent presence by category, roughly seven times as many appearances when the question is about function rather than brand [1][2][3].

The mechanism is unglamorous. A model's memory of names is built from training data, which is mostly the open web talking about you: Wikipedia, hundreds of G2 and Capterra reviews, a Crunchbase profile, Reddit threads, press. A product shipped last quarter has a homepage and maybe a Product Hunt launch, so there is almost nothing to recall [8]. Recall therefore tracks funding and press, both of which take years, which makes it a lagging indicator rather than a verdict on the product [7].

Three things stack on top. Training cutoff means anything published after the last run is invisible to memory until the next one [9]. Generic names actively hurt: Sierra, Wonder, Peek, Ray, Brew, Marx and Nora were all in the sample, and asked "what is Brew," the model has nothing to separate the email tool from the drink [10]. And a JavaScript-only site sends an empty body to crawlers including GPTBot and PerplexityBot, which do not execute JS, so the one source unambiguously about you says nothing [11]. The first two are slow work; the third is an afternoon, and it feeds the retrieval path that is already producing results [17].

Those results, as reported: Decagon, unknown by name, ranks first for AI customer service automation next to Ada, Intercom's Fin and Zendesk AI [12]. Lindy ranks first for AI automation platforms, ahead of Zapier and Make [13]. Sierra places fifth for customer experience platforms [14]. Half the sample was invisible by name and visible by function simultaneously, which is 24 of the 44 unrecognised products, or 55 percent of them [18][4].

Treat the numbers as one operator's single pass. The write-up does not name the models used, does not report repeat runs, and the category prompts were written by the tester rather than harvested from real buyers [19]. Search-grounded rankings are the volatile half of this, so a placement shown once is not a placement shown to be stable.

What to watch: whether fixing server-side rendering measurably changes category placement on the same prompts, whether the products in the 28 hold position across repeat runs, and whether any of the 44 crosses into name recall after the next training refresh [9][3][1].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories