Skip to content

Topic

LLM Inference Benchmarking

How tokens-per-second, tail latency, and time-to-first-token are measured and misread in local and served LLM benchmarks.

Current stories

build2 publishers

NVIDIA ships Groq 3 LPX and starts quoting inference in tokens per user, not per rack

The accelerator is in full production and the headline number is a single-request generation rate at 100,000 tokens of context. That is a different purchase order than throughput.

Perspective Coverage

3 publishers
Builder
Builder 48%
Operator
Operator 25%
Investor
Investor 27%

Reality

Evidence55
Adoption20
Hype gap+35
Incentives80
Confidence60