Build1 publisherNot yet confirmed elsewhere3 min readPublished
Exa hires a CTO for the part of search that has a power bill
Jeff Pinner arrives from Robinhood to run Exa's own index, GPU cluster and 323ms latency budget. Exa's stated targets imply roughly 97,000 requests in flight at once.
The Engineer · Build desk

What happened
- Will Bryk named Jeff Pinner as Exa's chief technology officer on August 22, hiring an engineer who had run technology at Robinhood and Lyft.
- The company trains its own retrieval models, operates its own web index and runs a GPU cluster to serve search under tight latency requirements.
- Exa's own figures put its crawlers at more than 500 billion URLs tracked and its user base at more than 400,000 developers.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost Keeping a full index fresh is a recurring bill that rivals aggregating other engines or answering over retrieved pages never pay, so part of the Series C is a subsidy for crawl coverage rather...
- decision Putting minute-long research jobs and millisecond API calls on the same accelerators forces a standing scheduling policy, which is a choice about whose work gets queued when the cluster fills.
- exposure Named customers building coding and sales products sit behind Exa's latency budget, so a tail regression surfaces inside their apps before it surfaces in Exa's own numbers.
- constraint Headroom for tens of thousands of simultaneous in-flight requests has to be paid for before the traffic exists, which limits how far the raise stretches once the index and the GPUs are both running.
Put Exa's two published numbers next to each other. The product page lists Instant mode at 323 milliseconds [11], and Bryk said the Series C would fund models and infrastructure able to process hundreds of thousands of searches per second [8]. Little's Law turns that pair into a capacity plan: at 300,000 requests per second, each holding a slot for 323 milliseconds, about 97,000 requests are in flight at any instant [20]. That figure, not the valuation, sets how many machines have to exist and how much index has to be resident in memory rather than fetched. An operator is hired to make it survivable, not to argue it down.
The harder part is that one estate has to carry two workloads of opposite shape. Exa is training retrieval models, operating its own web index and running a GPU cluster under tight latency requirements [4], and it also wants longer, more computationally expensive research jobs sitting alongside the fast API calls [5]. Those contend for the same accelerators. A research job that runs for minutes tolerates a queue; a 323ms Instant call does not [11]. Whoever writes the scheduler decides which customer degrades first, and that decision is permanent in a way a model release is not.
The funding history says how recent all of this is. Exa closed $250 million led by Andreessen Horowitz on May 20 at a reported $2.2 billion valuation [7], after $85 million led by Benchmark in September 2025 [9], against a disclosed total of at least $357 million [10]. Everything before the Series B therefore accounts for roughly $22 million [21], meaning about 94 percent of the disclosed capital arrived in the last two rounds [22]. The idea is five years old [2] and dates to a company founded in 2021 as Metaphor and renamed in January 2024 [18]; the money to operate at this size is months old.
Pinner's record fits the queueing problem more than the retrieval one. He spent almost a decade at Lyft on engineering infrastructure and the rideshare marketplace before serving as its CTO [12], then joined Robinhood as its first CTO in August 2024, after working as a distinguished engineer in Cruise's AI and robotics organization [13]. That Robinhood tenure ran about 21 months [23]; a regulatory filing shows the company separated him from the role on May 7, 2026, with benefits applicable to a termination without cause [3]. Runtimewire reads the appointment as pairing founder-led search research with an operator used to high-volume platforms [16]. Marketplaces and brokerages both punish tail latency in public, which is nearer to Exa's problem than anything about embeddings.
What a CTO hire does not fix is the shape of the demand. The crawlers tracking more than 500 billion URLs and the more than 400,000 developers are Exa's own figures [14], and at the Series C the company said more than 5,000 organizations used its services [15], which is 1.25 percent of that developer count and counts a different unit [24]. Named customers include Cursor, Cognition, HubSpot, OpenRouter and Monday.com [6]. Exa's case for the expense is that owning the index and the retrieval models lets it tune for agent-specific jobs such as source discovery for coding systems [17]. Rivals can crawl selected pages, aggregate existing search engines, or generate answers on top of retrieved sources [19], and none of those approaches pays to keep a full index fresh. That bill arrives every month, and it is now Pinner's to hold down.
What to watch
- Whether Exa publishes tail latency percentiles rather than a single Instant-mode figure, since that is what agent builders budget against.
- Whether the concurrency buildout arrives as owned accelerators or rented capacity, which the size and timing of any next raise would reveal.
- Whether the 500 billion URL and 400,000 developer counts are ever replaced with audited or customer-confirmed numbers.