Skip to content

BuildAlso reported elsewhere2 publishers3 min readPublished Updated

NVIDIA ships Groq 3 LPX and starts quoting inference in tokens per user, not per rack

The accelerator is in full production and the headline number is a single-request generation rate at 100,000 tokens of context. That is a different purchase order than throughput.

The Engineer · Build desk

How we use AISend a correction

What happened

  • NVIDIA said at Hot Chips that Groq 3 LPX, an interactive inference accelerator extending the Vera Rubin platform, is in full production.
  • Artificial Analysis recorded 3,400 output tokens per second on Gemma 4 31B at a 100,000-token context, which NVIDIA calls a record for that model.
  • Groq, the inference cloud whose name the part carries, is listed as a planned early adopter after Nebius.

Why it matters

  • decision Anyone sizing agent infrastructure now has to write a per-request generation floor at a named context length into the spec, because aggregate tokens per second no longer describes what the...
  • exposure With concurrency undisclosed, the sizing risk sits entirely with the buyer: a rate held at one request in flight and a rate held at fifty imply fleets that differ by an order of magnitude.
  • contradiction One NVIDIA release treats the Nebius adoption as done and the other as planned, so full production of the silicon says nothing yet about when the capacity is purchasable.
  • precedent If the accelerator stays hidden behind an unchanged API, hardware selection moves to the provider and inference contracts get argued in latency terms rather than part numbers.

Take the one number NVIDIA published and convert it into a unit of work. At 3,400 output tokens per second [3], a 10,000-token agent step lands in about 2.9 seconds [13]. The 4x framing [17] implies the unnamed nearest platform runs that step at roughly 850 tokens per second, or about 11.8 seconds [14]. Over a 200-step coding loop, that is under ten minutes against roughly 39 minutes [15].

Which is a real difference, and it is not the difference the same release advertises when it says coding tasks go from hours to minutes [18]. A 4x reduction applied to a two-hour task yields thirty minutes [16]. Tens of minutes is not minutes, so the two claims are measured against two different baselines, and neither baseline is named [5].

The specification change underneath is worth more than the benchmark. NVIDIA is defining interactivity as the token rate seen by an individual user, which sets how fast an agent closes each step of its loop [9]. Capacity has generally been bought as aggregate tokens per second per rack, or simply as accelerator count. A per-request floor at a stated context length is a different line item, and for a 100,000-token agent it is the line item that decides whether the thing is usable. The gap in the disclosure is concurrency: neither release says how many simultaneous requests were in flight when the 3,400 figure was recorded [5]. A box that holds that rate for one request and sags at fifty is a different machine from one that does not, and the published material cannot tell you which you are getting.

Then there is the awkward part for anyone hoping to specify this at all. Nebius CTO Danila Shtan says the acceleration arrives through the same API developers already use, with no migration to a new stack [7]. That is good for adoption and bad for procurement leverage: if LPX is invisible behind an existing endpoint, the buyer cannot order it, only ask for a latency number and let the provider choose the silicon. The two NVIDIA releases do not even agree on where Nebius is. One says Nebius is the first AI cloud to adopt the part [4]; the other says Nebius plans to bring it to its Token Factory inference platform [6]. Full production for the chip is not the same as a price and a date for the tokens.

The rest of the announcement reads as the same unbundling applied to the rack. NVIDIA says Vera Rubin now spans seven chips and five purpose-built racks [10]. CoreWeave is putting Spectrum-X Multiplane into production to link Vera Rubin racks across parallel switches [11], and SpaceXAI is taking Vera CPUs for the CPU-bound parts of agentic work, including orchestration, tool use and code execution [12]. The unit being sold is a rack matched to a phase of the agent loop. Capacity planning that still counts GPUs is measuring the wrong thing.

What to watch

  • Whether Artificial Analysis publishes the concurrency and batch conditions behind the 3,400 tokens per second run, and names the platform used for the 4x comparison.
  • A date and a price for Groq 3 LPX capacity inside Nebius Token Factory, which would turn the adoption claim into something a buyer can order.
  • How the arrangement behind the NVIDIA Groq 3 LPX name is explained when Groq, the inference cloud, brings the part into its own service.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence55
Adoption20
Hype gap+35
Incentives80
Confidence60

Perspective Coverage

3 publishers
Builder
Builder 48%
Operator
Operator 25%
Investor
Investor 27%
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    NVIDIA announced at Hot Chips that NVIDIA Groq 3 LPX, described as an interactive AI inference accelerator and an extension of the Vera Rubin platform, is now in full production.

  2. [2]

    The announcement was made at the Hot Chips conference in Palo Alto, California, with release material dated Tuesday, Aug. 24.

  3. [3]

    In Artificial Analysis benchmarking running Gemma 4 31B, an open source agentic model, with a 100,000-token context, Groq 3 LPX delivered a record 3,400 output tokens per second, which NVIDIA calls the fastest performance ever recorded for the model.

Sources

2 independent publishers whose own reporting we read for this story.

  1. developer.nvidia.com

    1 article · August 24, 2026

    How NVIDIA Groq 3 LPX Unlocks Ultrafast Interactivity at Long Context on NVIDIA Vera Rubin
  2. nvidianews.nvidia.com

    2 articles · August 24, 2026

    NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AI
  3. the-decoder.com

    2 articles · August 25, 2026

    Nvidia says its Groq 3 LPX is four times faster than Cerebras, but the math is more complicated

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories