BuildIndependently confirmed2 publishers3 min readPublished
Cerebras moves its product line from the wafer to the rack, and the CS-5 number carries a 2027 date
The Nexus rack, not the WSE-3, is now the thing Cerebras ships. That makes upgrade cadence the number buyers should price, and 10,000 tokens per second a target rather than a plan input.
The Engineer · Build desk

What happened
- Cerebras laid out the rack architecture behind CS-4 on August 25 at Hot Chips, a week after announcing the system, and previewed a CS-5 for 2027.
- The Nexus rack holds three WSE-3 Turbo processors, each in a rear-serviced compute backpack with its own power conversion, cooling, I/O and control electronics.
- The silicon is unchanged 5nm WSE-3; the performance doubling over CS-3 comes from higher clock speed paid for with more power and better cooling.
- It also claims up to 30 times GPU inference speed on a selected set of models, citing Artificial Analysis and its own August 2026 testing.
Why it matters
- decision Procurement questions move off the chip datasheet: what matters is whether each Nexus generation drops into cabinets already on the floor, which is a facilities and service-contract question.
- constraint Rear-access compute, front-side high voltage and up to 30 PSU modules per backpack mean a site that cannot support that service layout cannot take the upgrade cadence at any speed.
- exposure Anyone writing 10,000 tokens per second into a 2026 capacity plan is carrying the schedule and silicon risk of a system Cerebras has not yet described.
- contradiction Cerebras sells the rack as whole-system co-design while SemiAnalysis reads the networking gains as fairly small, so the buyer's own workload decides whether Nexus is an upgrade or a repackaging.
The power path is where the rack claim cashes out. Cerebras' Hot Chips write-up puts AC/DC conversion roughly 0.5 millimetres from the wafer, against about 50 millimetres in a conventional GPU system [9], which is a hundredfold cut in how far the last conversion stage sits from the load [18]. Taking the printed circuit board out of that final delivery path, Cerebras says, lets it push nearly twice as much power into the processor at almost the same voltage while limiting resistive loss [10]. That is the enabling move behind a clock-speed generation, and it belongs to the rack rather than to the silicon.
The arithmetic is worth doing slowly. The-decoder reports per-user throughput doubling on the same 5nm WSE-3 through clock speed bought with power and cooling [5], while wafer count per rack rose 50 percent [19]. If both hold, aggregate rack throughput is around three times CS-3, and the third wafer is buying concurrency rather than single-user latency [14]. Those are different purchases and they show up in different lines of a capacity model.
Against the 4,400 tokens per second per user Cerebras quotes for CS-4 [6], the 10,000-token CS-5 target needs a further 2.3 times [15] on hardware Cerebras has not described, in a year that is still a promise [8]. The nearest shipped datapoint runs well below both: an OpenAI GPT-5.6 Sol preview on Cerebras hardware at up to 750 output tokens per second in August, with access rationed while capacity expanded [11], or about 17 percent of the CS-4 headline [16]. Different model, different generation, but it is the figure a user actually saw, and OpenAI is also running Codex Spark on Cerebras hardware [7].
Memory is the constraint the token counts do not mention. Per-wafer capacity is unchanged at 44 GB [12], so a three-wafer rack holds 132 GB [20]. Cerebras also claims more than 1,000 tokens per second on models above 10 trillion parameters, from a mix of internal benchmarks, projections and extrapolations [23]. At even one byte per parameter, 10 trillion weights is 10 TB, or roughly 76 racks of wafer memory [17]. That headline is a fleet number wearing a system's clothes.
Disaggregated inference tells you who Cerebras thinks is signing. Prefill runs on someone else's silicon, with AMD Helios and AWS Trainium named as candidates, and Cerebras takes decode [4], so the sale is incremental decode capacity bolted onto a cluster the buyer already owns. Which is exactly why cadence is the number that matters. The 50 percent component reduction and the days-to-hours deployment claim come from Cerebras, not from field data [22]; the first customer install will settle it, because that claim gets tested in a change window rather than on a benchmark chart.
What to watch
- Whether Cerebras commits to CS-5 compute backpacks fitting Nexus racks already installed, which decides whether 2027 is a swap or a forklift.
- An independent production measurement of 4,400 tokens per second per user, rather than Cerebras' own August 2026 test set.
- Whether disaggregated prefill on AMD Helios or AWS Trainium ships with a named customer and a measured latency budget.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence56
- Adoption32
- Hype gap+34
- Incentives76
- Confidence62
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
CS-4 combines three WSE-3 Turbo processors inside Nexus, one per rear-mounted compute backpack, each containing its own power conversion, cooling, I/O and control electronics, with shared power equipment at the front of the rack so compute assemblies are installed or replaced from the rear.
- [2]
Each CS-4 backpack can house up to 30 PSU modules and connects to the rack's liquid-cooling network with leak detection and quick-disconnect valves, with high-voltage equipment at the front separated from water and compute at the rear.
- [3]
Cerebras detailed the rack architecture behind CS-4 on August 25 during the Hot Chips conference, a week after introducing the system, and previewed a CS-5 successor targeted for 2027.
- [4]
CS-4 supports disaggregated inference, with another type of processor handling prefill and Cerebras handling decode; Cerebras has named AMD Helios and AWS Trainium as possible prefill platforms, letting customers avoid replacing every part of an existing AI cluster.
- [5]
CS-4 still runs on the 5nm WSE-3 chip and doubles the performance of CS-3 by boosting clock speed through more power and better cooling.
- [6]
A single CS-4 rack holds three wafers instead of two and delivers up to 4,400 tokens per second per user.
- [7]
Cerebras hardware is used by OpenAI for Codex Spark, among others.
- [8]
Nexus makes Cerebras' rack, rather than a single wafer, the product roadmap, which could shorten upgrade cycles, but CS-5's 10,000-token target remains a 2027 promise.
- [9]
According to Cerebras' Hot Chips technical post, CS-4 places AC/DC conversion roughly 0.5 millimetres from the wafer, compared with about 50 millimetres in conventional GPU systems.
- [10]
Cerebras says removing the printed circuit board from the final delivery path lets CS-4 send nearly twice as much power to the processor at almost the same voltage while limiting resistive loss.
- [11]
RuntimeWire reported on August 13 that Cerebras hardware was powering an OpenAI GPT-5.6 Sol preview at up to 750 output tokens per second, with access limited while capacity expanded.
- [13]
Analysts at SemiAnalysis see the networking gains as fairly small.
- [14]
If per-user throughput doubles from clock speed while wafer count per rack rises 50 percent, aggregate rack throughput is about three times CS-3, meaning the added wafer buys concurrency rather than single-user speed.
- [15]
Reaching a 10,000 tokens per second per user target from CS-4's 4,400 requires a further 2.3 times improvement.
- [16]
The 750 output tokens per second measured on the August GPT-5.6 Sol preview is about 17 percent of CS-4's quoted 4,400 tokens per second per user.
- [17]
At one byte per parameter, a 10 trillion parameter model is 10 TB of weights, roughly 76 CS-4 racks' worth of on-wafer memory.
- [18]
Moving AC/DC conversion from about 50 millimetres to about 0.5 millimetres from the wafer is roughly a hundredfold reduction in that distance.
- [19]
Going from two wafers per rack to three is a 50 percent increase in wafer count per rack.
- [20]
A three-wafer CS-4 rack carries 132 GB of on-wafer memory.
- [21]
Cerebras calls CS-4 the industry's fastest AI accelerator and says it can deliver up to 30 times the inference speed of GPU systems across a selected set of models, attributing the comparisons to Artificial Analysis and its own August 2026 tests, with a disclaimer that results vary by model, workload, configuration and testing date.
- [22]
Cerebras says the backpack has 50% fewer components, uses 60% more automated manufacturing and can reduce deployment time from days to hours; those deployment figures come from Cerebras rather than published field data.
- [23]
Cerebras claims more than 1,000 tokens per second on models exceeding 10 trillion parameters, wafer-to-wafer latency as low as two microseconds, up to twice the performance of CS-3 and as much as 10 times the throughput per watt; some figures rely on internal benchmarks, projections or extrapolations rather than independent production comparisons.
Sources
2 independent publishers whose own reporting we read for this story.
- runtimewire.comCerebras details CS-4 rack and targets 10,000 tokens per second with CS-5
1 article · August 25, 2026
- the-decoder.comCerebras unveils CS-4 with double the performance on the same chip
1 article · August 24, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Data center power and coolingFollow
- AI Hardware Roadmaps and Upgrade CadenceFollow
- Rack-Scale AI SystemsFollow
- Inference Throughput BenchmarkingFollow
- Wafer-Scale ComputingFollow
- Disaggregated Prefill and DecodeFollow
Entities
- OpenAIFollow
- CerebrasFollow
- AWS TrainiumFollow
- Andrew FeldmanFollow
- Hot ChipsFollow
- AMD HeliosFollow
- Cerebras NexusFollow
- NvidiaFollow
- Cerebras CS-6Follow
- Cerebras CS-4Follow
- Artificial AnalysisFollow
- SeaMicroFollow
- WSE-3 / WSE-3 TurboFollow
- Codex SparkFollow
- AMDFollow
- Cerebras CS-3Follow
- ServeTheHomeFollow
- Cerebras CS-5Follow
- SemiAnalysisFollow
- GPT-5.6 SolFollow