Invest1 distinct publisher3 min readPublished
Ramp's Q4 2025 spend data sizes the model hosting and serving layer at $260m across roughly 1,900 buyers, which is 60 cents for every dollar those same companies hand straight to OpenAI and its closed-source peers.
The Investor · Invest desk

Compiled by The InvestorSomething wrong?How this is made
The ratio is the number to hold on to. If $260m is 60 per cent of what the same customers spent querying OpenAI, Anthropic, Google, xAI and Perplexity directly [3], the direct bill came to roughly $433m [1], the two lines together to about $693m, and the hosting-and-serving layer accounts for 37.5 per cent of everything that panel spent on running or calling models in the quarter [2]. That is a real business. It is also a much smaller thing than the phrase "AI infrastructure" usually carries, and Ramp itself describes the ecosystem as a niche player not yet putting pressure on closed-source vendors [2].
The narrowness is sharper on the buyer side. About 1,900 Ramp customers spend anything at all on this layer [4], which averages roughly $137,000 a quarter each, or about $547,000 annualised [4], spread very unevenly, presumably, since a market that runs from a serverless router like OpenRouter at one end to customer-operated bare-metal clusters at the other has no meaningful average account [7]. Ninety-four per cent of those buyers also pay a closed-source provider [5], which leaves something like 114 companies in the panel that host and serve without renting a frontier model at all [3]. That is the entire observable population of pure self-hosters.
And the two lines move together: more closed-source spend comes with more infrastructure spend, which Ramp reads as complementarity rather than substitution [6]. The mechanism is unglamorous. A team that has decided to put a model into production has to host it somewhere and make it callable by other services [10], and once that plumbing exists it reaches whatever the vendor's catalogue reaches, closed models included [9]. The price pressure that a credible open-weight alternative is supposed to exert on closed-source list prices [9] is not visible in this quarter's dollars [2].
Two readings I cannot test from the material. Corporate spend data catches pay-as-you-go and misses whatever sits inside committed multi-year compute contracts, so the $260m could be a floor and the 60 per cent understated; and a panel of companies that use a spend platform is not the same population as an insurer running models on hardware it already owns. Neither possibility is in the excerpt, and I would want the vendor-level breakdown before believing either.
This is probably wrong, but the view I would take is that NVIDIA, in licensing Groq's LPU designs to get into model hosting and serving [1], is buying an option on the customer count rather than on the current accounts. The falsifier is clean enough. If the next quarter shows roughly the same 1,900 buyers each spending materially more, then this is a whale market whose revenue concentration cuts both ways, and the interesting risk is not that a buyer cannot find a second provider but that a provider cannot survive losing four accounts. If the count moves instead, from 1,900 toward five figures at flat per-account spend, the narrowness was a timing artefact and the layer is doing what plumbing does. On present evidence, 1,900 buyers and 114 pure self-hosters [4][3] is a market of specialists, and specialists do not set prices for anyone else.
Ranked by verification strength, evidence, and original report placement.
Ramp finds the hosting and serving ecosystem is large and growing but still a niche player that is not yet putting pressure on closed-source companies.
Ramp customers spent $260M in Q4 2025 on AI infrastructure businesses, which is only 60% of what they spent on directly querying closed-source foundation model providers (OpenAI, Anthropic, Google, xAI, Perplexity).
Only a small subset of Ramp customers, about 1,900, spend on AI infrastructure.
94% of customers who spend on AI infrastructure also spend on closed-source models, so few customers spend only on self-hosting AI infrastructure.
Ramp separates the sector into three services: raw infrastructure bring-your-own-model self-run (examples CoreWeave and Nebius), hosted BYOM vendor-run inside a managed execution platform (examples Modal and Fireworks AI), and proxy/router serverless hosted models on a pay-as-you-go API (example OpenRouter).
For an LLM to be used in production it must be run somewhere (hosted) and be callable by other services (served), which Ramp calls vital for AI to break into the mainstream.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 30, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
invest
Nvidia's Perplexity talks move its money one layer further from its own chips1 distinct publisher
build
Hugging Face's $13B process puts most teams' model pipeline under a single owner2 distinct publishers
invest
The labs got better at watching their agents escape. They did not get better at stopping them.1 distinct publisher
build
Inco AI's DFlash 2: 21% longer accepted drafts for 1.3% latency and 18.5M parameters1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One ledger, no second look
Ramp is measuring spend it processes itself, which is about as close to a primary record as this market offers — and also a record no one else can audit. The quantitative spine holds together internally: the $260M total, the 60% ratio and the 94% overlap are mutually consistent, and Ramp volunteers that no standard definition of an LLM infrastructure company exists. What pulls the score down is everything the piece asserts without showing: the proportional co-movement between open and closed spend arrives with no distribution behind it, the promised growth rate never appears, and the NVIDIA–Groq deal that frames the whole report is stated in passing.
Real money, thin base
This is measured commercial adoption, not sentiment: money actually changed hands with named vendors, at scale, in a specific quarter. But the shape is narrow — about 1,900 buyers, near $137,000 each per quarter, and only some 114 companies using the layer without also paying a frontier API. That is a supplementary purchase pattern rather than a migration. The instrument also biases what adoption looks like: multi-year GPU commitments are invoiced, not swiped, so the self-run tier is likely undercounted while pay-as-you-go routers show up in full.
Cooler than its own headline
Ramp's numbers argue against the story most people tell about this layer. A market that has been narrated as the open-source counterweight to OpenAI turns out, in this panel, to be a niche that almost every buyer holds alongside a closed-source subscription. The findings are stated conservatively, hedged where they should be, and the one piece of promotional lift — the NVIDIA-and-Groq framing at the top, plus a 'fast-growing' headline with no growth rate attached — is the weakest-evidenced part of the piece. Understated overall, slightly oversold at the door.
The data is the marketing
Ramp sells spend management, and reports like this exist to establish it as the place where corporate AI spend gets counted first. That incentive rewards being cited, not reaching any particular verdict — which is partly why the conclusions run cautious and against the popular narrative. The subtler pull is definitional: Ramp decided which 25 companies constitute this market and which of three buckets each falls into, and every dollar figure inherits those choices along with the composition of its own customer base.
Right direction, unverifiable decimals
We would bet on the shape of this finding and not on its precision. The direction — hosting bought as a complement to frontier APIs by a small, high-spend cohort — is supported by the internal consistency of Ramp's figures and by how modestly it draws its own conclusions. Precision is another matter: one panel, one publisher, a quarter reported roughly eight months after it closed, an instrument blind to invoiced GPU contracts, and a framing news item nobody here confirms.