Invest1 publisher3 min readPublished
Rubin shows 7x tokens per megawatt and over 2x profit per gigawatt versus Blackwell, SemiAnalysis says
SemiAnalysis says its AgentX benchmark clocked Nvidia's Vera Rubin NVL72 at up to seven times Blackwell's token throughput per megawatt on pre-release software, and puts the profit gain at over two times per gigawatt.
The Investor · Invest desk

What happened
- SemiAnalysis published what it calls the first verified agentic inference results for Nvidia's Vera Rubin NVL72, measured on early pre-release software with its AgentX benchmark.
- On a 1.6-trillion-parameter DeepSeek model, those tests showed Rubin delivering up to seven times Blackwell's token throughput per megawatt.
- Jensen Huang's GTC 2026 graph put Vera Rubin NVL72 at three times Blackwell's performance per megawatt on models of one to three trillion parameters at around 200 tokens per second.
- SemiAnalysis estimates that even on early software builds Rubin can earn over two times more profit per gigawatt than the Blackwell platform.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- exposure Nvidia engineers, Huang among them, helped bring up the Rubin software and verify the results, so a buyer treating the seven-times figure as vendor-independent is relying on a methodology SemiAnalysis controls and Nvidia staff touched.
- precedent Two rounds of understated keynote numbers, 30 against 98 in 2024 and 3 against 7 now, teach large buyers to treat Nvidia's GTC per-watt figures as floors and to wait for third-party measurement before sizing an order.
- cost Blackwell owners spread server and networking capex over an assumed useful life, and a successor platform already earning twice the profit per gigawatt on immature software shortens the period over which that assumption holds.
- capability With TPUv7, Nvidia and AMD parts in one harness and SambaNova and Trainium coming, a procurement team can rank accelerators on replayed agentic traffic instead of on each vendor's chosen scenario.
Up to seven times the token throughput per megawatt. Over two times the profit per gigawatt. Divide 7 by 2 and you get 3.5, which is the most of that throughput gain that can go missing on the way to profit [18].
The reason sits in how SemiAnalysis builds the cost side. Its tokens-per-dollar figure normalizes total throughput by all-in serving cost, dollars per chip-hour times the number of chips used to serve [12]. Its hyperscaler ownership case builds that hourly cost from server and networking capex spread over an assumed useful life, plus colocation, power and cost of capital, on hyperscaler purchasing and financing terms [13]. Power is one of those inputs. For tokens per megawatt to rise up to sevenfold while profit per gigawatt rises a little past twice, either a Rubin megawatt costs more to own and operate than a Blackwell one, or revenue per token is lower, or both [19].
So a buyer sizing a build on performance per watt is pricing one input, and the return depends on the hourly cost of the hardware under those financing terms [13].
On the claims themselves, SemiAnalysis wrote that "Jensen needs to stop sandbagging his performance claims at GTC" [5]. It cites the precedent. At GTC 2024 Huang said GB200 NVL72 would deliver 30 times Hopper's performance, and SemiAnalysis measured 98 times, about 3.3 times the claim [6][17]. The 2026 gap is narrower, 7 measured against 3 presented, about 2.3 times [3][4][16].
The measurement was made with Nvidia in the room. SemiAnalysis thanks Huang, Ian Buck, Nick Comly, Kedar Potdar, Rohit Nagraj and the Mainland China TensorRTLLM Team for helping on next-gen Rubin software bring-up and helping verify the agentic benchmark results [11]. It also says the benchmark has been widely reproduced, validated and/or supported by almost every major buyer of compute, naming Google Cloud, Microsoft Azure, Oracle and Meta, along with vLLM, SGLang, PyTorch, Huggingface and labs including OpenAI, Qwen and Moonshot Kimi [10]. That wording puts reproduction and endorsement in the same sentence, so the count of parties that re-ran AgentX cannot be separated from the count that approves of it.
AgentX replays real-world agentic traffic across a fleet of thousands of chips, sessions of tens or hundreds of turns in which the ratio of cached input to uncached tends toward 1 as the turn count grows [9][15].
In my view the over-2x profit per gigawatt is the figure that goes into a purchase model, and the 7x is a measurement of one line in it. The counter-argument is easy to state. At a site where the grid interconnect is the binding limit, megawatts are the scarce input and capital is not, and seven times the tokens out of the same substation is worth a higher hourly cost. A third path is that inference prices fall as Rubin supply arrives, which compresses profit per gigawatt on both platforms and leaves the throughput ratio where it is.
SemiAnalysis expects the gap to widen as the Rubin software stack and kernel libraries mature and developers gain experience optimizing for it [8]. If the mature-stack figure moves toward three or four times the profit per gigawatt while Rubin's hourly cost stays near Blackwell's, then performance per watt was the right frame and this piece was wrong. The other thing that would undo it is the traffic assumption: a provider whose sessions recompute more prefill than AgentX's cached-input ratio implies gets a different ranking [15].
What to watch
- AMD's MI455X UALoE72 numbers on AgentX, which AMD has committed to collaborating on, and where they land against Rubin on tokens per dollar.