buildOne report1 publisher Cerebras says splitting inference stages across chip types gave 5x more throughput from the same number of its systems without slowing token generation. Because the count covers only Cerebras hardware, the figure does not yet show what a mixed-chip fleet costs per unit of work.
Reality
- Evidence30
- Adoption20
- Hype gap+30
- Incentives75
- Confidence40
buildOne report1 publisher d-Matrix CTO Sudeep Bhoja says Raptor's 1,000 tokens per second per user comes from a simulated 72-card system, with full racks not due until Q4 2027. Whole-system speed, power and cost per request are still unmeasured, so buyers are working from early silicon tests and simulations.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+35
- Incentives70
- Confidence40
buildOne report1 publisher Gimlet Labs plans 100 MW of Cerebras-powered inference capacity, pairing wafer-scale chips with GPUs for different phases of each request. For buyers, the quoted 3,000 tokens per second is a company target until measured results appear.
Reality
- Evidence40
- Adoption15
- Hype gap+40
- Incentives70
- Confidence40
Groq 3 LPX is in full production with Nebius as the named first customer. The benchmark is one model at one context length, and the 4x claim does not quite get from hours to minutes.
Reality
- Evidence45
- Adoption25
- Hype gap+35
- Incentives70
- Confidence55
buildConfirmed2 publishers The Nexus rack, not the WSE-3, is now the thing Cerebras ships. That makes upgrade cadence the number buyers should price, and 10,000 tokens per second a target rather than a plan input.
Reality
- Evidence56
- Adoption32
- Hype gap+34
- Incentives76
- Confidence62