Skip to content

model

Gemma 4 31B

Model measured at 9.81s warm average and 10.41 tok/s, slower than the 120B-A12B Nemotron build despite far fewer parameters; followed the format.

Current stories

build2 publishers

NVIDIA ships Groq 3 LPX and starts quoting inference in tokens per user, not per rack

The accelerator is in full production and the headline number is a single-request generation rate at 100,000 tokens of context. That is a different purchase order than throughput.

Perspective Coverage

3 publishers
Builder
Builder 48%
Operator
Operator 25%
Investor
Investor 27%

Reality

Evidence55
Adoption20
Hype gap+35
Incentives80
Confidence60