Skip to content

Product1 publisher3 min readPublished

CoreWeave rents out Vera Rubin NVL72 racks backed by one customer's 4.8x throughput claim

CoreWeave now rents Nvidia's Vera Rubin NVL72 racks, with first customer Cognition reporting up to 4.8x the token throughput on SWE-2 inference. That figure is a best case from a single inference workload, and it is the only customer result behind the launch.

The Product Desk · Product desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying CoreWeave rents out Vera Rubin NVL72 racks backed by one customer's 4.8x throughput claim
Generated illustration

What happened

  • CoreWeave says Cognition began using the Vera Rubin systems in early September, a couple of months after CoreWeave claimed the first fully working NVL72 rack.
  • Nvidia claims Rubin delivers five times the inference and 3.5 times the training performance of Blackwell, the GPU generation it replaces.
  • CoreWeave also plans to rent Nvidia's Vera CPU on its own as bare metal, with 128 CPUs and 11,264 cores in a single rack.
  • At the same Fully Connected conference in San Francisco, home-care company Ennoble Care signed on for RTX Pro 6000 Blackwell nodes on CoreWeave's Kubernetes service.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • exposure Capacity plans sized on Cognition's 4.8x are sized on a best case, so a team whose own gain lands lower will need more racks than it budgeted for.
  • constraint Because the NVL72 is fully liquid-cooled, only data halls built for liquid can host it, so Rubin capacity will be limited to providers that have them.
  • capability A GPU cloud renting bare-metal CPUs lets agent builders run sandboxes and model inference with one provider once Vera moves past early access.

Cognition was seeing performance gains within days of getting onto the new racks, according to CoreWeave. "Bringing up Nvidia Vera Rubin NVL72 so quickly, and having a customer already seeing performance gains within days, is the payoff from years of engineering our platform across GPU generations," said Chen Goldberg, CoreWeave's executive vice president of product and engineering [5]. With customers like Cognition, he said, that investment "shows up in the ability to get production workloads running within days" [6].

The performance figures are Nvidia's. The company claims gains over Blackwell in both inference and training [8], and it says the NVL72's cable-free trays cut rack installation from two hours to five minutes [9]. That is a 24-fold reduction [1]. In public, one customer has run one inference workload, and Cognition's Silas Alberti put the gain at up to 4.8x in total token throughput on SWE-2 [7]. That sits just under Nvidia's 5x inference claim [8]. "Up to" makes it the best result Cognition reported.

Total throughput fits the work Goldberg described. "When it comes to agentic tasks, long contexts, repeated model calls, and thousands of concurrent tasks put pressure on the entire platform," he said [10]. A fleet of coding agents needs tokens across the whole rack. A developer waiting on one answer needs low latency, and the reported figure covers throughput only [7].

The report does not include a price or the hardware Cognition's 4.8x was measured against. A finance lead asking for cost per token will need both.

Ennoble Care, the other customer announced at the conference, chose the older RTX Pro 6000 Blackwell nodes [14]. Its CTO, Jonathan Taylor, said of CoreWeave: "It gives us reserved capacity we can count on, a Kubernetes environment our team can move into quickly, and engineers who answer the phone" [15].

The CPU side is at an earlier stage. CoreWeave has traditionally sold access to GPUs, and it now plans to rent Vera as bare metal [11]. Its Vera racks work out to 88 cores per CPU [2]. According to Corey Sanders, CoreWeave's SVP of product, the offering is in its early stages. He said some customers are expected to start testing in the coming weeks, and a few already have early access [12].

A team deciding on this has two things to check. The first is whether its workload resembles Cognition's, meaning throughput-bound inference with many concurrent calls [10]. The second is whether CoreWeave will quote a price and let the team run its own job before it commits. When both hold, the number to compute is the team's own measured throughput gain divided by the price premium over its current hardware. That ratio has to clear 1 by enough to pay for the migration. By CoreWeave's account, Cognition's migration took days [6]. For a team whose workload matches but who has no quote yet, the 4.8x is a good reason to request one. Training jobs and latency-sensitive jobs have only Nvidia's own projections to plan against [8].

What to watch

  • CoreWeave publishing per-hour or per-token pricing for Vera Rubin NVL72, which would allow a cost-per-token comparison against Blackwell.
  • A second customer, or Cognition itself, reporting typical gains or any training result that can be checked against Nvidia's 3.5x claim.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories