g factor's Qwen 3.8 27B benchmark has Together AI fastest at one stream, at 189.61 tok/s, while four of five engines finish within about 10% at 64 streams. Choosing a provider from these numbers starts with knowing how many streams the deployment will run at once.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+18
- Incentives72
- Confidence45
The accelerator is in full production and the headline number is a single-request generation rate at 100,000 tokens of context. That is a different purchase order than throughput.
Perspective Coverage
3 publishers
- Builder
- Builder 48%
- Operator
- Operator 25%
- Investor
- Investor 27%
Reality
- Evidence55
- Adoption20
- Hype gap+35
- Incentives80
- Confidence60
The designated successor to GenAI-Perf splits load generation from record processing so the benchmark client stops hitting Python's GIL. Teams holding a GenAI-Perf baseline inherit a port and a re-run.
Reality
- Evidence55
- Adoption20
- Hype gap+15
- Incentives75
- Confidence55
Speculative decoding speedups depend on the data, and most published ones come from high-level scripts on narrow datasets. SPEED-Bench's authors argue the honest measurement happens inside vLLM or TensorRT-LLM, across concurrencies.
Reality
- Evidence42
- Adoption21
- Hype gap+18
- Incentives58
- Confidence44
A dev.to explainer on local inference benchmarks makes a point worth pinning up: tokens per second is a function of how many users you tested with, not a property of the hardware.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+14
- Incentives42
- Confidence41