OpenAI has contracted 750MW of Cerebras wafer-scale inference capacity, delivered in three 250MW segments by the end of 2028. The dates are contractual, but neither company has published pricing, latency figures or a plan for which customers get access.
Reality
- Evidence45
- Adoption25
- Hype gap+10
- Incentives45
- Confidence40
Gimlet Labs plans 100 MW of Cerebras-powered inference capacity, pairing wafer-scale chips with GPUs for different phases of each request. For buyers, the quoted 3,000 tokens per second is a company target until measured results appear.
Reality
- Evidence40
- Adoption15
- Hype gap+40
- Incentives70
- Confidence40
More memory on a Cerebras wafer arrives with CS-6, two generations out. Everything shipping before then, Nexus included, is rack engineering around the SRAM budget the wafer already has, and long contexts are where that budget hurts.
Reality
- Evidence42
- Adoption33
- Hype gap+34
- Incentives72
- Confidence41
The first third-party benchmark of the LP30 rack came in at roughly four times the next-fastest public endpoint, measured one request at a time on a model small enough to fit.
Reality
- Evidence58
- Adoption20
- Hype gap+32
- Incentives74
- Confidence55
The Nexus rack, not the WSE-3, is now the thing Cerebras ships. That makes upgrade cadence the number buyers should price, and 10,000 tokens per second a target rather than a plan input.
Reality
- Evidence56
- Adoption32
- Hype gap+34
- Incentives76
- Confidence62
The first multi-wafer Cerebras system pairs a doubled clock with rebuilt power delivery and interconnect. The compute claims rest on WSE-3 Turbo dies that are otherwise unchanged.
Reality
- Evidence54
- Adoption34
- Hype gap+34
- Incentives76
- Confidence61