Build2 publishers2 min readPublished
Cognition reports 4.8x token throughput on Nvidia Vera Rubin in its own coding-agent test
Cognition says its SWE-2 model produced up to 4.8 times more total token throughput on Nvidia's Vera Rubin NVL72 than on GB200, in a test on CoreWeave. The company-run result points to more agent capacity per system, but whether long coding tasks get cheaper depends on what the new systems cost per hour.
The Engineer · Build desk

What happened
- CoreWeave identified Cognition as the first customer running production workloads on Vera Rubin, after the two stood up its cluster in early September.
- The test used SWE-2 inference workloads built from a sample of tasks in Cognition's FrontierCode benchmark, measured against a GB200 NVL72 baseline.
- CoreWeave separately reported 3.8 times higher output-token throughput on reinforcement-learning workloads.
- The September 30 announcement followed Cognition's September 8 Series E, which it said raised more than $2 billion at a $48 billion valuation.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Cognition has to choose whether the extra capacity buys more concurrent sessions, shorter agent cycles or heavier use of each deployment, and capacity pricing decides which of those pays off.
- exposure Devin training, reinforcement learning and production inference already run on thousands of CoreWeave GPUs, so how fast Cognition moves onto Vera Rubin depends on one provider's rollout.
- precedent Vera Rubin's first production customer is a coding-agent company, so the first customer benchmark buyers can point to reflects a coding agent's token mix and not their own.
'Up to' is the most honest phrase in the announcement. RuntimeWire calls the 4.8x an upper-end throughput claim for one SWE-2 test [13]. It describes the figures as vendor-associated results, not an independent replication [11].
I think the most useful line in the release is CoreWeave's statement that the gain came with no loss in generation speed [7]. A throughput gain bought by slowing each request would cost a coding agent time on every step. The agent makes repeated model calls as it plans, edits, tests and revises code [5]. CoreWeave's reinforcement-learning multiple counts output tokens only, so it measures something different from the total-token 4.8x [8][1].
A coding agent that runs through many steps can consume far more inference than a single prompt-and-answer exchange [6]. The business paying for that inference is also growing. By the company's own measure, Cognition's run-rate revenue rose from $492 million in May to almost $900 million [17], roughly 1.8 times [1].
For the multiple to carry over to another operator, several things would need to hold. That operator's token mix would need to resemble SWE-2 on Cognition's task sample [9]. Nvidia describes those tasks as real-world software-engineering work [10]. Its GB200 baseline would need to match the one Cognition's engineers measured against [9]. And its workload would need to land near the top of the range [13]. The published materials do not include absolute throughput, system pricing, energy use or enough configuration detail to reproduce the comparison [12]. None of those conditions can be checked.
Cost per token is a system's hourly price divided by the tokens it produces in that hour [2]. At the full 4.8x, a Vera Rubin NVL72 matches a GB200 NVL72 on cost per token if it costs 4.8 times as much per hour, and beats it at any lower price [2]. At a smaller multiple, the break-even price falls in step [2]. Token counts also stop short of completed work. By RuntimeWire's reading, the result does not establish 4.8 times more completed coding tasks or lower total cost across workloads [18]. Whether an agent's output stays useful as throughput rises is a separate question [14].
What to watch
- CoreWeave's hourly pricing for Vera Rubin NVL72 against GB200 NVL72, which sets whether 4.8x becomes a lower cost per token.
- An independent benchmark of Vera Rubin against GB200 on agent workloads that publishes absolute throughput and full configuration.
- Whether Cognition reports completed tasks or cost per task on Vera Rubin, in place of token throughput alone.