Skip to content

Product1 publisher3 min readPublished

LlamaIndex builds document models on rented GPUs that have to absorb its customers' paperwork spikes

LlamaIndex runs a workload that is 75% inference on rented GPUs, parsing millions of document pages a day that arrive in bursts. Its CEO's account at the Fully Connected event describes a buyer for whom guaranteed capacity matters more than owning hardware.

The Product Desk · Product desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying LlamaIndex builds document models on rented GPUs that have to absorb its customers' paperwork spikes
Photo: siliconangle.com

What happened

  • LlamaIndex started as an open-source framework for retrieval-augmented generation and has turned into a model builder, according to co-founder and CEO Jerry Liu.
  • CoreWeave used Fully Connected to launch Forge, a development layer that runs training, inference, evaluation and agent development in one connected environment.
  • CoreWeave's Lukas Biewald said the company follows Nvidia-recommended standard networking protocols, while AWS leans on proprietary APIs that make workloads hard to move.
  • Analyst Dave Vellante said CoreWeave customers are pushing for shorter contracts, spot pricing and on-demand access, all priced above long-term take-or-pay deals.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • cost Startups with spiky inference loads face the highest prices, since the short and on-demand terms that fit their demand cost more than long commitments.
  • decision A team that keeps its web service on AWS or GCP now has to decide whether sending only inference to a second provider is worth running two clouds.
  • constraint A latency gain tied to how one provider networks and powers its chips has to be tested on a team's own model before it can justify a contract.
  • precedent If Liu is right that the post-training loop gets automated for everyone, more small teams become GPU renters with inference-heavy loads, the customers CoreWeave's wider stack is meant to keep.

A year ago LlamaIndex's compute footprint barely existed [5]. Now its finance, legal and insurance customers send paperwork in bursts, and the company parses it on capacity it rents [4][1]. Jerry Liu, its co-founder and chief executive, described the requirement in one sentence. "We really, really need to make sure that we have the right capacity to serve our customers without getting throttled," Liu said [6].

In that workload, inference is three times the size of training [1]. The models are LlamaIndex's own. "We post-train open-weight models, we gather our own datasets and we make it really, really good at analyzing and reading documents to basically extract that data," Liu said [2]. He also put latency on the list, next to cost. "We care a lot about making sure that we can actually tailor everything we're doing at the Pareto frontier of performance, cost, and latency for our customers," he said [18].

CoreWeave's pitch at the event was wider than any of that. The company is building out networking, storage and software as inference demand grows faster than training [12]. Lukas Biewald, its senior vice president of AI initiatives, joined when CoreWeave bought Weights & Biases, the startup he co-founded [8]. He said he came from software doubting that chip configuration mattered much [8]. Biewald said performance can differ by orders of magnitude depending on how the chips are networked and how power is distributed [9]. The coverage does not include a measurement behind that claim.

His openness argument is about living alongside the big clouds. "Everyone's going to host their web service on AWS or GCP, not on CoreWeave. CoreWeave is okay with that, so CoreWeave plays much more nicely with the other clouds," Biewald said [11].

Liu is the only customer quoted in either report. The other voices are CoreWeave's own executive and theCUBE's analysts, and theCUBE is a paid media partner for Fully Connected whose coverage CoreWeave sponsored [15].

The buying behavior shows up in contract terms. "The prices for those types of structures are much, much higher, and people are willing to pay," Dave Vellante, chief analyst at theCUBE Research, said of short-term and on-demand deals [14]. The customers Vellante describes are paying higher prices for shorter commitments [13].

For a team signing a GPU contract, the decision turns on the share of its compute that is inference and how far a peak day sits above an ordinary one. A mostly-training team on a steady schedule belongs on a long take-or-pay contract, the cheaper structure [13], and can treat the wider software stack as optional. A mostly-inference team with spikes is in LlamaIndex's position [3][4]. It is buying protection from throttling. The premium for short or on-demand terms is the price of that protection, to be set against what a delayed customer batch costs. Steady inference can commit for its base load and keep spot capacity for overflow. Spiky training can rent by the run and keep commitments small.

What to watch

  • CoreWeave publishing latency comparisons across networking and power setups that buyers could check against Biewald's orders-of-magnitude claim.
  • Any CoreWeave disclosure of its contract mix showing whether the move toward short-term and spot deals Vellante described is happening at scale.
  • Whether LlamaIndex keeps renting as its inference volume grows, or signs long commitments for its base load.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories