ProductIndependently confirmed2 publishers3 min readPublished Updated
Nvidia puts a decode chip in the rack, and inference planning gets a second variable
Groq 3 LPX is in full production with Nebius as the named first customer. The benchmark is one model at one context length, and the 4x claim does not quite get from hours to minutes.
The Product Desk
What happened
- Nvidia said at Hot Chips 2026 that Groq 3 LPX, a dedicated inference accelerator, has entered full production.
- It splits the inference job inside the rack: Vera Rubin GPUs take context ingestion, the new parts take token generation.
- Artificial Analysis clocked it at 3,400 tokens per second on the open-source Gemma 4 31B model with a 100,000-token context window.
Why it matters
- decision Buying inference capacity now means choosing a ratio between context and generation silicon per workload, and the wrong guess leaves one half of the rack idle.
- cost The licence bill and the two hires are sunk cost that has to come back through what decode capacity sells for, and no price has been put on it.
- constraint Anyone sizing an agent latency budget on this is extrapolating from a single measured point, which is thinner evidence than a capacity commitment usually rests on.
- precedent With generation speed sold as its own line item, every competing inference platform will now be asked for its decode part and its number on the same test.
Nvidia's own comparison puts the part four times ahead of rival platforms on latency-sensitive work [15], which sets the baseline it beat at roughly 850 tokens per second [16]. That is a checkable number. The other claim in the same announcement, that multistep agentic tasks finish in minutes instead of hours [13], is not: a 4x factor applied to a three-hour job leaves 45 minutes [14]. Either the tasks in question sat close to an hour rather than six, or the saving comes from something other than decode throughput. The 4x is the figure worth holding Nvidia to.
The second gap is that the source never says whether the measured rate is one stream or the whole rack, and that distinction is the product. Per stream it works out to about 0.29 milliseconds a token, so a 10,000-token reasoning chain clears in roughly three seconds [18], which is the number an agent loop budget is built from. Spread across the accelerators a full rack can hold, the same figure would be about 13 tokens per second each [17], and nobody would publish that as a record. Read it as single-stream, then ask what the batch size and the concurrency were.
Then the planning problem, which is the part that lands on operators this quarter. The rack pairs dozens of Vera Rubin GPUs doing context ingestion with a much larger population of decode parts [5][6], and neither the intended ratio for a given workload mix nor whether a buyer can vary it is stated. An agent that reads a 100,000-token context and emits 200 tokens has a different balance than one that plans in long chains from a short prompt. Guess wrong and you have idle silicon on one side of the interconnect. Huang's phrasing, workload-optimized AI factory configurations for the era of agentic AI [9], is honest about what is being sold: a configuration, not a chip. Configurations have to be sized, and the sizing evidence here is one model at one context length [7]. Worth noting too that the announcement's own text calls the rack's accelerators LP30 while the product is Groq 3 LPX [11], so treat the configuration detail as provisional.
The clearest fact in the release is the $20 billion paid in December for Groq's technology, along with founder Jonathan Ross and president Sunny Madra [8]. Nvidia decided that decode-specialised silicon was worth buying rather than waiting to build, and took the people who had already shipped it. Nebius CTO Danila Shtan makes the same argument from the buyer's side, that generation is the phase that decides how responsive a system feels [3].
What is absent: price, power draw, and how much of this Nebius is installing [10]. Nebius says it will run the chips in its Token Factory platform [2]. Until it publishes a rate per million output tokens, the responsiveness claim is a specification rather than an operating cost.
What to watch
- Whether Artificial Analysis publishes run conditions for the 3,400 tokens per second figure: single stream or aggregate, batch size, and which platform the 4x comparison was against.
- Whether Nebius states how many LPX racks are live in Token Factory and what it charges per million output tokens against GPU-only inference.
- Whether the next named customer can buy decode capacity separately, or whether it only ships inside a Vera Rubin NVL72 configuration.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence45
- Adoption25
- Hype gap+35
- Incentives70
- Confidence55
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Nvidia said its dedicated AI inference accelerator Groq 3 LPX has entered full production, announced at Hot Chips 2026.
ReportedSupportedSource: Nvidia, reported by SiliconANGLE3 sources— create a free account to open themView cited source - [2]
Nvidia said neocloud provider Nebius Group N.V. is the first customer to commit to the chip, and Nebius said it will use it in Nebius Token Factory, its production inference platform.
- [3]
Nebius CTO Danila Shtan said generation is the phase of inference that determines how responsive an AI system actually is, and that Nebius is the first AI cloud to bring the chip to production via Nebius Token Factory.
ReportedSupportedSource: Danila Shtan, Nebius CTO3 sources— create a free account to open themView cited source - [4]
Groq 3 LPX is described by Nvidia as a purpose-built extension to its Vera Rubin data center platform, designed to deliver ultra-fast token generation for responsive agentic AI workloads.
- [5]
The design disaggregates context processing from token generation: the Vera Rubin NVL72 rack-scale platform, powered by dozens of Vera Rubin GPUs, handles large-scale context ingestion, while decode workloads are offloaded to the new accelerators.
- [6]
According to Nvidia, a full rack-scale deployment can harness up to 256 accelerators, linked by its chip interconnects, working in tandem with the GPUs as a unified inference engine.
- [7]
In benchmark tests by Artificial Analysis, Groq 3 LPX output 3,400 tokens per second running the open-source Gemma 4 31B agentic model with a 100,000-token context window.
ReportedSupportedSource: Artificial Analysis, cited by Nvidia2 sources— create a free account to open themView cited source - [8]
The chip was built using technology licensed from Groq Inc.; Nvidia paid the startup $20 billion in December for access and hired its founder Jonathan Ross and President Sunny Madra as part of the deal.
- [9]
Jensen Huang said Vera Rubin extends Nvidia's inference work with workload-optimized AI factory configurations designed for the era of agentic AI, and that LPX advances the performance frontier for ultra-fast token generation.
ReportedSupportedSource: Jensen Huang, Nvidia CEO2 sources— create a free account to open themView cited source - [10]
The announcement as reported does not disclose a price, a power figure, or how many racks or accelerators Nebius will deploy.
- [11]
The announcement text refers to a rack harnessing up to 256 'LP30' accelerators while the product elsewhere is named Groq 3 LPX.
- [12]
Agentic workloads can crunch thousands of tokens across chains of reasoning, tool calls and code execution, which the source says sometimes produces massive decode latency and user-visible delays.
- [13]
Nvidia says multistep agentic tasks can be completed in minutes instead of hours, and that the chip eliminates the tradeoff between throughput and response times.
- [14]
A 4x speedup applied to a three-hour task leaves 45 minutes, so 'minutes instead of hours' only holds for tasks near the bottom of the hours range.
- [15]
Nvidia says the chip is four times more responsive for latency-sensitive workloads compared with rival platforms.
- [16]
Nvidia's 4x responsiveness claim implies the comparison platforms it beat were producing about 850 tokens per second under the same conditions.
- [17]
If 3,400 tokens per second were an aggregate figure across a fully populated rack, it would be about 13 tokens per second per accelerator.
- [18]
At 3,400 tokens per second per stream, one token takes about 0.29 milliseconds and a 10,000-token generation completes in about 2.9 seconds.
Sources
2 independent publishers whose own reporting we read for this story.
- datacenterdynamics.comNvidia’s ultra-low-latency AI inference LPX racks hit full production
1 article · August 25, 2026
- siliconangle.comNvidia’s dedicated inference accelerator Groq 3 LPX enters full production to supercharge AI agents
2 articles · August 25, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
Entities
- Sinclair SchullerFollow
- Groq 3 LPXFollow
- Danila ShtanFollow
- Hot ChipsFollow
- NvidiaFollow
- Jensen HuangFollow
- Artificial AnalysisFollow
- Nebius Token FactoryFollow
- Sunny MadraFollow
- Gemma 4 31BFollow
- Jonathan RossFollow
- Vera RubinFollow
- GroqFollow
- Dell Technologies Inc.Follow
- NebiusFollow