Product1 distinct publisher3 min readUpdated
Groq 3 LPX is in full production with Nebius as the named first customer. The benchmark is one model at one context length, and the 4x claim does not quite get from hours to minutes.
The Product Desk · Product desk
Compiled by The Product DeskSomething wrong?How this is made
Nvidia's own comparison puts the part four times ahead of rival platforms on latency-sensitive work [7], which sets the baseline it beat at roughly 850 tokens per second [12]. That is a checkable number. The other claim in the same announcement, that multistep agentic tasks finish in minutes instead of hours [8], is not: a 4x factor applied to a three-hour job leaves 45 minutes [13]. Either the tasks in question sat close to an hour rather than six, or the saving comes from something other than decode throughput. The 4x is the figure worth holding Nvidia to.
The second gap is that the source never says whether the measured rate is one stream or the whole rack, and that distinction is the product. Per stream it works out to about 0.29 milliseconds a token, so a 10,000-token reasoning chain clears in roughly three seconds [15], which is the number an agent loop budget is built from. Spread across the accelerators a full rack can hold, the same figure would be about 13 tokens per second each [14], and nobody would publish that as a record. Read it as single-stream, then ask what the batch size and the concurrency were.
Then the planning problem, which is the part that lands on operators this quarter. The rack pairs dozens of Vera Rubin GPUs doing context ingestion with a much larger population of decode parts [4][5], and neither the intended ratio for a given workload mix nor whether a buyer can vary it is stated. An agent that reads a 100,000-token context and emits 200 tokens has a different balance than one that plans in long chains from a short prompt. Guess wrong and you have idle silicon on one side of the interconnect. Huang's phrasing, workload-optimized AI factory configurations for the era of agentic AI [10], is honest about what is being sold: a configuration, not a chip. Configurations have to be sized, and the sizing evidence here is one model at one context length [6]. Worth noting too that the announcement's own text calls the rack's accelerators LP30 while the product is Groq 3 LPX [17], so treat the configuration detail as provisional.
The clearest fact in the release is the $20 billion paid in December for Groq's technology, along with founder Jonathan Ross and president Sunny Madra [9]. Nvidia decided that decode-specialised silicon was worth buying rather than waiting to build, and took the people who had already shipped it. Nebius CTO Danila Shtan makes the same argument from the buyer's side, that generation is the phase that decides how responsive a system feels [11].
What is absent: price, power draw, and how much of this Nebius is installing [16]. Nebius says it will run the chips in its Token Factory platform [3]. Until it publishes a rate per million output tokens, the responsiveness claim is a specification rather than an operating cost.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Nvidia said its dedicated AI inference accelerator Groq 3 LPX has entered full production, announced at Hot Chips 2026.
Groq 3 LPX is described by Nvidia as a purpose-built extension to its Vera Rubin data center platform, designed to deliver ultra-fast token generation for responsive agentic AI workloads.
Nvidia said neocloud provider Nebius Group N.V. is the first customer to commit to the chip, and Nebius said it will use it in Nebius Token Factory, its production inference platform.
The design disaggregates context processing from token generation: the Vera Rubin NVL72 rack-scale platform, powered by dozens of Vera Rubin GPUs, handles large-scale context ingestion, while decode workloads are offloaded to the new accelerators.
According to Nvidia, a full rack-scale deployment can harness up to 256 accelerators, linked by its chip interconnects, working in tandem with the GPUs as a unified inference engine.
In benchmark tests by Artificial Analysis, Groq 3 LPX output 3,400 tokens per second running the open-source Gemma 4 31B agentic model with a 100,000-token context window.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Vendor announcement with one narrow third-party datapoint
Everything traces to a single trade report of a single vendor announcement. There is one third-party measurement (Artificial Analysis) but it covers one model at one context length and is reported without stating whether the throughput is per stream or per rack. Core planning facts — price, power, deployment scale — are absent, and the announcement text is internally inconsistent about the part name.
Full production claimed, one named customer, volumes unknown
Adoption evidence is real but thin: the chip is said to be in full production, one neocloud (Nebius) has committed to run it in a named production platform, and a broader Vera Rubin platform commitment from SpaceX is disclosed. No shipment counts, rack counts, revenue, or general availability details are given, and no customer is described as already serving traffic on the accelerator.
Performance framing runs ahead of the arithmetic
The vendor language — 'eliminates the tradeoff between throughput and response times', 'minutes instead of hours', '4x more responsive' — is stronger than the single benchmark and its own arithmetic support. A 4x speedup leaves a three-hour task at 45 minutes, the rival baseline of roughly 850 tokens per second is never named or measured, and the headline throughput figure lacks a unit basis. The underlying architectural change is substantive, which keeps the gap moderate rather than extreme.
Launch-day vendor and customer promotion
Every performance and positioning statement originates with parties that gain from it: Nvidia announcing its own SKU on stage, Nvidia needing to show product from a $20 billion Groq licensing deal and founder hires, and Nebius marketing itself as the first AI cloud to production. The reporting reproduces those quotations without adversarial testing, and the third-party benchmark is invoked by the vendor as validation.
Single-publisher, single-announcement basis
The factual spine — production status, first customer, rack topology, benchmark number, licensing deal — is consistently reported and internally coherent, so low-level facts are reasonably firm. But there is only one publisher and one announcement behind everything, no independent confirmation of the benchmark's conditions, and an unresolved naming and unit ambiguity, so confidence in the performance conclusions stays modest.
build
NVIDIA ships Groq 3 LPX and starts quoting inference in tokens per user, not per rack1 distinct publisher
product
Nvidia circles Rebellions because the low-power inference tier is not optional2 distinct publishers
invest
The 81% Problem: AI's Star CEOs Are Polling Badly With The People They Need To Hire1 distinct publisher
product
Nvidia's $105bn guarantee, not OpenAI's balance sheet, is what makes Ohio's 8 gigawatts buildable1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.