Skip to content

Product2 publishers2 min readPublished Updated

A 100-billion-document index for rent, and retrieval turns into a thing labs buy

Keenable left stealth with $26 million and says AI labs already query its index during training and at runtime. The scarcity argument comes from the investor who led the round.

The Product Desk · Product desk

Photograph accompanying A 100-billion-document index for rent, and retrieval turns into a thing labs buy
Photo: techcrunch.com

What happened

  • Keenable emerged from stealth with $26 million in seed funding led by Accel, alongside Conviction Partners and angel investors.
  • The company says its API, built on an index of more than 100 billion documents, is already in production at several AI labs and inference providers, in training and at runtime.
  • It was founded by Andrey Styskin, formerly head of Yandex's search, AI and cloud division, and Matthias Petri, his colleague from Amazon search infrastructure work.
  • Keenable has a partnership with voice AI company Gradium for live information retrieval, and will not name its other customers.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • decision Crawling the web yourself now has to be justified against renting a corpus whose scanning cost is spread across other tenants, a build-or-buy call most labs previously settled by default.
  • constraint If the incumbents keep withdrawing API access, agent builders are left choosing between bundled terms set by a giant and a seed-stage index they cannot audit.
  • exposure Anyone whose inference path already runs through a rented index inherits its outages, freshness gaps and pricing changes, and buyers cannot benchmark against peers who stay unnamed.
  • precedent Shipping a query language rather than a raw index is the move that converts retrieval from a swappable input into a rewrite cost, which is how this layer stops being commoditised.

Web-scale retrieval has one economic property that decides the rest: the expensive part is scanning, not answering. Styskin puts it as a technical problem, saying that without index structures fine-tuned for a specific task the cost of serving and scanning the whole internet is enormous, and that the work is in narrowing the search space fast [4]. Read as a business statement instead, it says a crawl is a large fixed cost looking for more tenants. One index queried by several labs at training time and again at inference is the same corpus billed twice, which is the only way the arithmetic gets tolerable for anyone who is not Google.

The scarcity half of the argument comes from Accel's Zhenya Loginov, who led the investment and says AI players have very few options at web scale, particularly as Google and Microsoft move to close their existing search APIs rather than cannibalise themselves, favouring bundled deals with selected partners [6]. That is the load-bearing claim in the whole story, and it is made by the person who wrote the cheque. Keenable's own supporting evidence is similarly indirect: Styskin says he saw Cloudflare data showing AI crawlers taking a growing share of search volume [7], and the company will not name the labs and inference providers it says run its API in production [5].

Then there is a number the announcement leaves lying around. Twenty-six million dollars against more than 100 billion documents works out to roughly 26 cents per thousand documents already indexed [11], and Styskin's only comment on what the index cost was "Don't ask - it is painfully expensive" [10]. The round plainly did not pay for the crawl. It is going to headcount: 15 engineers across the U.S. and Europe, doubling by the end of the year, and the stated purpose is a go-to-market motion [9], which is roughly 30 people [12]. A company that already has the asset and is now hiring sellers is a company that thinks the buyers exist.

What would make that dependency sticky is the next product rather than the current one. WebQueryLanguage is meant to let AI systems assemble an answer from several web sources when no single one contains it [8]. An index is substitutable. A query language that agent code is written against is not, and Brave, Exa and a Google busy rebuilding its own search experience are all competing on the substitutable part [13].

What to watch

  • Whether any AI lab or inference provider confirms on the record that it buys retrieval from Keenable, since the customer list is currently the company's word.
  • Any further move by Google or Microsoft on search API access, which would strengthen or undercut the scarcity argument behind the round.
  • Whether WebQueryLanguage ships with freshness and pricing numbers attached, or stays a description of an upcoming product.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories