BuildAlso reported elsewhere2 publishers3 min readPublished Updated
NVIDIA ships Groq 3 LPX and starts quoting inference in tokens per user, not per rack
The accelerator is in full production and the headline number is a single-request generation rate at 100,000 tokens of context. That is a different purchase order than throughput.
The Engineer · Build desk
What happened
- NVIDIA said at Hot Chips that Groq 3 LPX, an interactive inference accelerator extending the Vera Rubin platform, is in full production.
- Artificial Analysis recorded 3,400 output tokens per second on Gemma 4 31B at a 100,000-token context, which NVIDIA calls a record for that model.
- Groq, the inference cloud whose name the part carries, is listed as a planned early adopter after Nebius.
Why it matters
- decision Anyone sizing agent infrastructure now has to write a per-request generation floor at a named context length into the spec, because aggregate tokens per second no longer describes what the...
- exposure With concurrency undisclosed, the sizing risk sits entirely with the buyer: a rate held at one request in flight and a rate held at fifty imply fleets that differ by an order of magnitude.
- contradiction One NVIDIA release treats the Nebius adoption as done and the other as planned, so full production of the silicon says nothing yet about when the capacity is purchasable.
- precedent If the accelerator stays hidden behind an unchanged API, hardware selection moves to the provider and inference contracts get argued in latency terms rather than part numbers.
Take the one number NVIDIA published and convert it into a unit of work. At 3,400 output tokens per second [3], a 10,000-token agent step lands in about 2.9 seconds [13]. The 4x framing [17] implies the unnamed nearest platform runs that step at roughly 850 tokens per second, or about 11.8 seconds [14]. Over a 200-step coding loop, that is under ten minutes against roughly 39 minutes [15].
Which is a real difference, and it is not the difference the same release advertises when it says coding tasks go from hours to minutes [18]. A 4x reduction applied to a two-hour task yields thirty minutes [16]. Tens of minutes is not minutes, so the two claims are measured against two different baselines, and neither baseline is named [5].
The specification change underneath is worth more than the benchmark. NVIDIA is defining interactivity as the token rate seen by an individual user, which sets how fast an agent closes each step of its loop [9]. Capacity has generally been bought as aggregate tokens per second per rack, or simply as accelerator count. A per-request floor at a stated context length is a different line item, and for a 100,000-token agent it is the line item that decides whether the thing is usable. The gap in the disclosure is concurrency: neither release says how many simultaneous requests were in flight when the 3,400 figure was recorded [5]. A box that holds that rate for one request and sags at fifty is a different machine from one that does not, and the published material cannot tell you which you are getting.
Then there is the awkward part for anyone hoping to specify this at all. Nebius CTO Danila Shtan says the acceleration arrives through the same API developers already use, with no migration to a new stack [7]. That is good for adoption and bad for procurement leverage: if LPX is invisible behind an existing endpoint, the buyer cannot order it, only ask for a latency number and let the provider choose the silicon. The two NVIDIA releases do not even agree on where Nebius is. One says Nebius is the first AI cloud to adopt the part [4]; the other says Nebius plans to bring it to its Token Factory inference platform [6]. Full production for the chip is not the same as a price and a date for the tokens.
The rest of the announcement reads as the same unbundling applied to the rack. NVIDIA says Vera Rubin now spans seven chips and five purpose-built racks [10]. CoreWeave is putting Spectrum-X Multiplane into production to link Vera Rubin racks across parallel switches [11], and SpaceXAI is taking Vera CPUs for the CPU-bound parts of agentic work, including orchestration, tool use and code execution [12]. The unit being sold is a rack matched to a phase of the agent loop. Capacity planning that still counts GPUs is measuring the wrong thing.
What to watch
- Whether Artificial Analysis publishes the concurrency and batch conditions behind the 3,400 tokens per second run, and names the platform used for the 4x comparison.
- A date and a price for Groq 3 LPX capacity inside Nebius Token Factory, which would turn the adoption claim into something a buyer can order.
- How the arrangement behind the NVIDIA Groq 3 LPX name is explained when Groq, the inference cloud, brings the part into its own service.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence55
- Adoption20
- Hype gap+35
- Incentives80
- Confidence60
Perspective Coverage
3 publishers- Builder
- Builder 48%
- Operator
- Operator 25%
- Investor
- Investor 27%
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
NVIDIA announced at Hot Chips that NVIDIA Groq 3 LPX, described as an interactive AI inference accelerator and an extension of the Vera Rubin platform, is now in full production.
- [2]
The announcement was made at the Hot Chips conference in Palo Alto, California, with release material dated Tuesday, Aug. 24.
- [3]
In Artificial Analysis benchmarking running Gemma 4 31B, an open source agentic model, with a 100,000-token context, Groq 3 LPX delivered a record 3,400 output tokens per second, which NVIDIA calls the fastest performance ever recorded for the model.
- [4]
NVIDIA states that Nebius is the first AI cloud to adopt NVIDIA Groq 3 LPX.
- [5]
Neither NVIDIA release states the number of concurrent requests or batch size at which the 3,400 output tokens per second figure was measured, and neither names the nearest alternative platform used for the 4x comparison.
- [6]
NVIDIA states that Nebius plans to bring NVIDIA Groq 3 LPX to Nebius Token Factory, its production inference platform.
- [7]
Nebius chief technology officer Danila Shtan said generation is the phase of inference that determines how responsive an AI system actually is, and that Nebius is making every step of an agent's loop feel instant through the same API developers are already using, with no migration to a new stack.
ReportedSupportedSource: Danila Shtan, CTO, Nebius3 sources— create a free account to open themView cited source - [8]
NVIDIA says that following Nebius, the purpose-built AI inference cloud Groq plans to be among the platform's earliest adopters.
- [9]
NVIDIA says Groq 3 LPX is purpose-built to extend Vera Rubin's interactivity, which it defines as the rate at which tokens are generated for an individual user, determining how quickly an agent can complete each step of its work.
- [10]
NVIDIA describes Vera Rubin as its most extensive AI factory platform, built through extreme codesign across seven chips and five purpose-built racks.
- [11]
CoreWeave is deploying Spectrum-X Multiplane in production, connecting NVIDIA Vera Rubin racks using multiple parallel switches.
- [12]
SpaceXAI plans to deploy NVIDIA Vera CPUs to accelerate CPU-intensive agentic AI work including orchestration, tool use, code execution, data processing and simulation, and plans to build its future AI architecture around Vera Rubin from terrestrial data centres to orbital satellites.
- [13]
At 3,400 output tokens per second, a 10,000-token generation step takes about 2.9 seconds.
- [14]
If Groq 3 LPX is 4x the nearest alternative at 3,400 tokens per second, the implied alternative rate is about 850 tokens per second, or about 11.8 seconds for a 10,000-token step.
- [15]
A 200-step agent loop of 10,000-token steps takes about 9.8 minutes at 3,400 tokens per second and about 39 minutes at the implied 850 tokens per second.
- [16]
A 4x speedup applied to a two-hour task produces a thirty-minute task, so a claim of hours becoming minutes requires a baseline slower than the nearest alternative platform.
- [17]
NVIDIA says Groq 3 LPX provides 4x faster responsiveness for agents and latency-sensitive workloads than the nearest alternative platform.
- [18]
NVIDIA says Groq 3 LPX enables agentic tasks such as coding in minutes versus hours.
Sources
2 independent publishers whose own reporting we read for this story.
- developer.nvidia.comHow NVIDIA Groq 3 LPX Unlocks Ultrafast Interactivity at Long Context on NVIDIA Vera Rubin
1 article · August 24, 2026
- nvidianews.nvidia.comNVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AI
2 articles · August 24, 2026
- the-decoder.comNvidia says its Groq 3 LPX is four times faster than Cerebras, but the math is more complicated
2 articles · August 25, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.