Skip to content

ProductIndependently confirmed2 publishers3 min readPublished Updated

Nvidia puts a decode chip in the rack, and inference planning gets a second variable

Groq 3 LPX is in full production with Nebius as the named first customer. The benchmark is one model at one context length, and the 4x claim does not quite get from hours to minutes.

The Product Desk

How we use AISend a correction

What happened

  • Nvidia said at Hot Chips 2026 that Groq 3 LPX, a dedicated inference accelerator, has entered full production.
  • It splits the inference job inside the rack: Vera Rubin GPUs take context ingestion, the new parts take token generation.
  • Artificial Analysis clocked it at 3,400 tokens per second on the open-source Gemma 4 31B model with a 100,000-token context window.

Why it matters

  • decision Buying inference capacity now means choosing a ratio between context and generation silicon per workload, and the wrong guess leaves one half of the rack idle.
  • cost The licence bill and the two hires are sunk cost that has to come back through what decode capacity sells for, and no price has been put on it.
  • constraint Anyone sizing an agent latency budget on this is extrapolating from a single measured point, which is thinner evidence than a capacity commitment usually rests on.
  • precedent With generation speed sold as its own line item, every competing inference platform will now be asked for its decode part and its number on the same test.

Nvidia's own comparison puts the part four times ahead of rival platforms on latency-sensitive work [15], which sets the baseline it beat at roughly 850 tokens per second [16]. That is a checkable number. The other claim in the same announcement, that multistep agentic tasks finish in minutes instead of hours [13], is not: a 4x factor applied to a three-hour job leaves 45 minutes [14]. Either the tasks in question sat close to an hour rather than six, or the saving comes from something other than decode throughput. The 4x is the figure worth holding Nvidia to.

The second gap is that the source never says whether the measured rate is one stream or the whole rack, and that distinction is the product. Per stream it works out to about 0.29 milliseconds a token, so a 10,000-token reasoning chain clears in roughly three seconds [18], which is the number an agent loop budget is built from. Spread across the accelerators a full rack can hold, the same figure would be about 13 tokens per second each [17], and nobody would publish that as a record. Read it as single-stream, then ask what the batch size and the concurrency were.

Then the planning problem, which is the part that lands on operators this quarter. The rack pairs dozens of Vera Rubin GPUs doing context ingestion with a much larger population of decode parts [5][6], and neither the intended ratio for a given workload mix nor whether a buyer can vary it is stated. An agent that reads a 100,000-token context and emits 200 tokens has a different balance than one that plans in long chains from a short prompt. Guess wrong and you have idle silicon on one side of the interconnect. Huang's phrasing, workload-optimized AI factory configurations for the era of agentic AI [9], is honest about what is being sold: a configuration, not a chip. Configurations have to be sized, and the sizing evidence here is one model at one context length [7]. Worth noting too that the announcement's own text calls the rack's accelerators LP30 while the product is Groq 3 LPX [11], so treat the configuration detail as provisional.

The clearest fact in the release is the $20 billion paid in December for Groq's technology, along with founder Jonathan Ross and president Sunny Madra [8]. Nvidia decided that decode-specialised silicon was worth buying rather than waiting to build, and took the people who had already shipped it. Nebius CTO Danila Shtan makes the same argument from the buyer's side, that generation is the phase that decides how responsive a system feels [3].

What is absent: price, power draw, and how much of this Nebius is installing [10]. Nebius says it will run the chips in its Token Factory platform [2]. Until it publishes a rate per million output tokens, the responsiveness claim is a specification rather than an operating cost.

What to watch

  • Whether Artificial Analysis publishes run conditions for the 3,400 tokens per second figure: single stream or aggregate, batch size, and which platform the 4x comparison was against.
  • Whether Nebius states how many LPX racks are live in Token Factory and what it charges per million output tokens against GPU-only inference.
  • Whether the next named customer can buy decode capacity separately, or whether it only ships inside a Vera Rubin NVL72 configuration.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence45
Adoption25
Hype gap+35
Incentives70
Confidence55
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Nvidia said its dedicated AI inference accelerator Groq 3 LPX has entered full production, announced at Hot Chips 2026.

    ReportedSupportedSource: Nvidia, reported by SiliconANGLE3 sources— create a free account to open themView cited source
  2. [2]

    Nvidia said neocloud provider Nebius Group N.V. is the first customer to commit to the chip, and Nebius said it will use it in Nebius Token Factory, its production inference platform.

  3. [3]

    Nebius CTO Danila Shtan said generation is the phase of inference that determines how responsive an AI system actually is, and that Nebius is the first AI cloud to bring the chip to production via Nebius Token Factory.

    ReportedSupportedSource: Danila Shtan, Nebius CTO3 sources— create a free account to open themView cited source

Sources

2 independent publishers whose own reporting we read for this story.

  1. datacenterdynamics.com

    1 article · August 25, 2026

    Nvidia’s ultra-low-latency AI inference LPX racks hit full production
  2. siliconangle.com

    2 articles · August 25, 2026

    Nvidia’s dedicated inference accelerator Groq 3 LPX enters full production to supercharge AI agents

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories