Build1 distinct publisher3 min readPublished
The first third-party benchmark of the LP30 rack came in at roughly four times the next-fastest public endpoint, measured one request at a time on a model small enough to fit.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The memory ledger is the part to read first. A full rack is 256 LP30s, which totals 128GB of on-die SRAM feeding 40 PB/s of aggregate bandwidth against 315 PFLOPS of FP8 compute, with 350 ns of chip-to-chip latency in an MGX liquid-cooled, Vera Rubin-compatible frame [8]. One Rubin GPU carries 288GB of HBM4 [11]. The rack that posted the record therefore holds about 44 percent of the memory of a single GPU [3], and every byte of a served model has to live inside it.
Nvidia's own arithmetic puts a 31-billion-parameter model at FP8 at around 62 LPUs just to hold weights [11]. That is about 31GB, roughly a quarter of the rack, leaving on the order of 97GB for the 100K-context state and everything else [4]. Gemma 4 31B is dense and fits in one rack [7]. A large mixture-of-experts model runs to four figures of chips across several racks [11], and that case did not come up on stage [7].
The measurement deserves the same close reading. Artificial Analysis took the median of 50 sequential client requests at a concurrency of one, against a private pre-release endpoint served through Google Cloud [5]. Serving one request at a time is the condition that yields the highest per-user token rate a machine can post, and it does not map onto the multi-tenant load the rival endpoints were carrying [c5b]. Nvidia's own demo figure of 10,996 tokens per second on the same model is 3.2 times the audited number [6][2]; Arsovski labelled it self-reported before telling the room the aim was "third-party verified independent benchmarks that you guys can trust" [c6b].
Underneath is a pipeline with no caches, no branch prediction and no out-of-order execution, scheduled by the compiler at clock-cycle granularity, with weights resident in SRAM rather than streamed from HBM [9]. Determinism then earns its keep twice over. Because power draw is known cycle by cycle, Nvidia pre-orders current from the rack's regulators, which it says cuts voltage droop by more than 60 percent and overshoot by more than 70 percent against an uncompensated load [13]. Per-block scheduling also equalizes heat instead of throttling to the hottest tile, worth roughly 10 to 11 percent more performance under a fixed thermal limit, according to Arsovski [14]. Across racks, chips run against a single virtual clock and each one routes as well as computes, so the fabric needs no adaptive routing or congestion sensing [16].
That puts weight on the failure answer. Asked in Q&A about a chip dying mid-workload, Arsovski said users "would experience the exact same as any other hardware in the industry" and would "just checkpoint it or reconfigure the hardware" [15]. On a rack where one model's weights are spread over dozens of chips [11] and the network has nothing sensing congestion to steer around a hole [16], reconfiguration is not a small operation.
Ranked by verification strength, evidence, and original report placement.
Igor Arsovski, Groq's former chief architect and now Nvidia's VP of hardware, presented the Groq 3 LPX rack architecture at Hot Chips 2026.
Artificial Analysis measured the Groq 3 LPX rack at 3,431 output tokens per second on a 100K-context Gemma 4 31B reasoning workload, roughly four times the 870 tokens per second of the next-fastest public endpoint; this was the first third-party benchmark of the hardware.
Arsovski said the rack is already in production, built on the LP30 chip Nvidia obtained through its $20 billion Groq deal in December 2025.
The same Groq deal pushed the Rubin CPX, which the LPX rack replaces, off Nvidia's roadmap.
Artificial Analysis ran the comparison on a private, pre-release Gemma 4 31B endpoint served through Google Cloud, taking the median of 50 sequential client requests at a concurrency of one, while the public providers it measured against ran shared production serverless endpoints.
Serving one request at a time produces the highest per-user token rate the hardware can post, and it is not directly comparable to the multi-tenant conditions the other endpoints run under.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One disclosed third-party run plus detailed vendor architecture claims
The architecture and performance detail is specific and attributed, and the benchmark comes from an outside firm with its methodology stated, which is stronger than a pure vendor deck. But it is a single measurement of a single dense model on a private pre-release endpoint at concurrency one, reported by one publisher, with the vendor's own figure 3.2x higher and self-flagged; power, thermal and split-inference gains are Nvidia-measured.
Vendor-asserted production, no external deployment evidence
The only adoption signals are Nvidia's own statement that the rack is in production and one benchmarking firm's access to a private pre-release endpoint. No customer, volume, pricing or public endpoint is disclosed, and the design is positioned as a co-processor attached to Vera Rubin rather than as a standalone deployed platform.
Headline speedup rests on the most favourable measurement conditions
The 'roughly four times' framing compares a single-stream, concurrency-one run on a private pre-release endpoint against shared multi-tenant production endpoints — a comparison the reporting itself calls not directly comparable — on a dense model chosen to fit one rack, while trillion-parameter MoE behaviour, multi-tenant throughput and cost go unaddressed and the vendor's louder 10,996 tokens/sec number is self-reported. The overstatement is in the framing rather than fabricated data: the caveats, capacity math and self-reported flag are all disclosed in the same article, which keeps the gap moderate.
Acquirer validating a $20B deal, vendor-supplied benchmark endpoint, competing vendor in the same session
Nvidia is presenting silicon it bought for $20 billion and for which it cancelled Rubin CPX, with the acquired company's former chief architect as presenter — a strong incentive to show a decisive number. The third-party benchmark ran on a private, pre-release endpoint the vendor side made available, and Cerebras used the same session to push its own unverified wafer-scale comparisons, while the publisher promoted free access to its Hot Chips coverage window.
Detailed and self-caveating, but single-publisher and single-run
Confidence is supported by direct quotes, named attribution and disclosed methodology, and by the article surfacing its own limitations. It is capped by there being one publisher, one benchmark run, one model, no independent replication, and no corroboration of the production claim or the Nvidia-measured system-level gains.
invest
Nvidia's $20bn Groq buy becomes shipping racks, and the tape reads it as execution risk2 distinct publishers
product
Cerebras's CS-4 is three old wafers in a new rack: price the packaging, not the silicon2 distinct publishers
build
Cerebras moves its product line from the wafer to the rack, and the CS-5 number carries a 2027 date2 distinct publishers
product
Nvidia circles Rebellions because the low-power inference tier is not optional2 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.