Gimlet Labs plans 100 MW of Cerebras-powered inference capacity, pairing wafer-scale chips with GPUs for different phases of each request. For buyers, the quoted 3,000 tokens per second is a company target until measured results appear.
Reality
- Evidence40
- Adoption15
- Hype gap+40
- Incentives70
- Confidence40
More memory on a Cerebras wafer arrives with CS-6, two generations out. Everything shipping before then, Nexus included, is rack engineering around the SRAM budget the wafer already has, and long contexts are where that budget hurts.
Reality
- Evidence42
- Adoption33
- Hype gap+34
- Incentives72
- Confidence41
The Nexus rack, not the WSE-3, is now the thing Cerebras ships. That makes upgrade cadence the number buyers should price, and 10,000 tokens per second a target rather than a plan input.
Reality
- Evidence56
- Adoption32
- Hype gap+34
- Incentives76
- Confidence62
The Cambridge company says routing agentic work across Nvidia, AMD and Google chips doubles accuracy at a quarter of the cost. No customer has confirmed those numbers.
Reality
- Evidence42
- Adoption22
- Hype gap+38
- Incentives74
- Confidence55
The first multi-wafer Cerebras system pairs a doubled clock with rebuilt power delivery and interconnect. The compute claims rest on WSE-3 Turbo dies that are otherwise unchanged.
Reality
- Evidence54
- Adoption34
- Hype gap+34
- Incentives76
- Confidence61