Published · yesterdayBuild3 min read
NVIDIA ships Groq 3 LPX and starts quoting inference in tokens per user, not per rack
The accelerator is in full production and the headline number is a single-request generation rate at 100,000 tokens of context. That is a different purchase order than throughput.
Written for builders.See today for builders
What happened
- NVIDIA said at Hot Chips that Groq 3 LPX, an interactive inference accelerator extending the Vera Rubin platform, is in full production.
- Artificial Analysis recorded 3,400 output tokens per second on Gemma 4 31B at a 100,000-token context, which NVIDIA calls a record for that model.
- Groq, the inference cloud whose name the part carries, is listed as a planned early adopter after Nebius.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decisionAnyone sizing agent infrastructure now has to write a per-request generation floor at a named context length into the spec, because aggregate tokens per second no longer describes what the...
- exposureWith concurrency undisclosed, the sizing risk sits entirely with the buyer: a rate held at one request in flight and a rate held at fifty imply fleets that differ by an order of magnitude.
- contradiction
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
NVIDIA announced at Hot Chips that NVIDIA Groq 3 LPX, described as an interactive AI inference accelerator and an extension of the Vera Rubin platform, is now in full production.
ReportedView cited source - [2]
The announcement was made at the Hot Chips conference in Palo Alto, California, with release material dated Tuesday, Aug. 24.
ReportedView cited source - [3]
In Artificial Analysis benchmarking running Gemma 4 31B, an open source agentic model, with a 100,000-token context, Groq 3 LPX delivered a record 3,400 output tokens per second, which NVIDIA calls the fastest performance ever recorded for the model.
ReportedView cited source - [4]
NVIDIA says Groq 3 LPX provides 4x faster responsiveness for agents and latency-sensitive workloads than the nearest alternative platform.
ReportedView cited source - [5]
NVIDIA says Groq 3 LPX is purpose-built to extend Vera Rubin's interactivity, which it defines as the rate at which tokens are generated for an individual user, determining how quickly an agent can complete each step of its work.
ReportedView cited source - [6]
NVIDIA says Groq 3 LPX enables agentic tasks such as coding in minutes versus hours.
ReportedView cited source
Sources & coverage · 3 publishers
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- developer.nvidia.comTanya LenzyesterdayHow NVIDIA Groq 3 LPX Unlocks Ultrafast Interactivity at Long Context on NVIDIA Vera Rubin
- nvidianews.nvidia.comyesterdayNVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AI
- nvidianews.nvidia.com


