Build2 distinct publishers3 min readPublished
The Nexus rack, not the WSE-3, is now the thing Cerebras ships. That makes upgrade cadence the number buyers should price, and 10,000 tokens per second a target rather than a plan input.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The power path is where the rack claim cashes out. Cerebras' Hot Chips write-up puts AC/DC conversion roughly 0.5 millimetres from the wafer, against about 50 millimetres in a conventional GPU system [5], which is a hundredfold cut in how far the last conversion stage sits from the load [17]. Taking the printed circuit board out of that final delivery path, Cerebras says, lets it push nearly twice as much power into the processor at almost the same voltage while limiting resistive loss [6]. That is the enabling move behind a clock-speed generation, and it belongs to the rack rather than to the silicon.
The arithmetic is worth doing slowly. The-decoder reports per-user throughput doubling on the same 5nm WSE-3 through clock speed bought with power and cooling [12], while wafer count per rack rose 50 percent [18]. If both hold, aggregate rack throughput is around three times CS-3, and the third wafer is buying concurrency rather than single-user latency [19]. Those are different purchases and they show up in different lines of a capacity model.
Against the 4,400 tokens per second per user Cerebras quotes for CS-4 [13], the 10,000-token CS-5 target needs a further 2.3 times [20] on hardware Cerebras has not described, in a year that is still a promise [2]. The nearest shipped datapoint runs well below both: an OpenAI GPT-5.6 Sol preview on Cerebras hardware at up to 750 output tokens per second in August, with access rationed while capacity expanded [11], or about 17 percent of the CS-4 headline [21]. Different model, different generation, but it is the figure a user actually saw, and OpenAI is also running Codex Spark on Cerebras hardware [16].
Memory is the constraint the token counts do not mention. Per-wafer capacity is unchanged at 44 GB [14], so a three-wafer rack holds 132 GB [22]. Cerebras also claims more than 1,000 tokens per second on models above 10 trillion parameters, from a mix of internal benchmarks, projections and extrapolations [10]. At even one byte per parameter, 10 trillion weights is 10 TB, or roughly 76 racks of wafer memory [23]. That headline is a fleet number wearing a system's clothes.
Disaggregated inference tells you who Cerebras thinks is signing. Prefill runs on someone else's silicon, with AMD Helios and AWS Trainium named as candidates, and Cerebras takes decode [8], so the sale is incremental decode capacity bolted onto a cluster the buyer already owns. Which is exactly why cadence is the number that matters. The 50 percent component reduction and the days-to-hours deployment claim come from Cerebras, not from field data [4]; the first customer install will settle it, because that claim gets tested in a change window rather than on a benchmark chart.
Ranked by verification strength, evidence, and original report placement.
CS-4 combines three WSE-3 Turbo processors inside Nexus, one per rear-mounted compute backpack, each containing its own power conversion, cooling, I/O and control electronics, with shared power equipment at the front of the rack so compute assemblies are installed or replaced from the rear.
Each CS-4 backpack can house up to 30 PSU modules and connects to the rack's liquid-cooling network with leak detection and quick-disconnect valves, with high-voltage equipment at the front separated from water and compute at the rear.
Cerebras detailed the rack architecture behind CS-4 on August 25 during the Hot Chips conference, a week after introducing the system, and previewed a CS-5 successor targeted for 2027.
CS-4 supports disaggregated inference, with another type of processor handling prefill and Cerebras handling decode; Cerebras has named AMD Helios and AWS Trainium as possible prefill platforms, letting customers avoid replacing every part of an existing AI cluster.
CS-4 still runs on the 5nm WSE-3 chip and doubles the performance of CS-3 by boosting clock speed through more power and better cooling.
A single CS-4 rack holds three wafers instead of two and delivers up to 4,400 tokens per second per user.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed architecture, vendor-only performance
The structural facts are well specified and consistent across two publishers: three WSE-3 Turbo wafers per Nexus rack, rear-serviceable backpacks with local power and cooling, up to 30 PSU modules, 0.5 mm AC/DC conversion, unchanged 44 GB per wafer, 4,400 tokens/second per user. Nearly all quantitative performance evidence, however, traces to Cerebras' own announcement, Hot Chips post and August 2026 tests; the two independent datapoints are one reported 750 tokens/second preview observation and SemiAnalysis' modest-networking read. No field deployment data supports the days-to-hours or throughput-per-watt figures.
Announced product, prior-generation usage only
CS-4 is announced and architecturally disclosed but no customer deployment, shipment or pricing evidence appears. The demand signals in the cluster attach to existing Cerebras hardware rather than CS-4: OpenAI's Codex Spark usage and a GPT-5.6 Sol preview that ran at up to 750 output tokens/second with rationed access while capacity expanded. Prefill partners (AMD Helios, AWS Trainium) are named as possible, not contracted.
Headline multiples run ahead of measurement
The marketing frame - 'industry's fastest', up to 30x GPU inference speed, 10x throughput per watt, 10,000 tokens/second - sits well above what is measured. The one observed serving rate is 750 tokens/second, roughly 17 percent of the quoted 4,400; the 10,000 figure is a 2027 forward-looking statement needing a further 2.3x; the uplift comes from clocking the same 5nm WSE-3 harder rather than new silicon; and the only outside analyst view calls networking gains fairly small. The rack-engineering substance (power conversion proximity, serviceability, prefill/decode disaggregation) is real and comparatively under-emphasized, which caps the gap short of pure vapor.
Vendor-run launch narrative with market stakes
Every core performance number originates with the seller at its own launch and conference slot, benchmarked partly by itself, with comparisons chosen across a 'selected set of models' and forward-looking targets placed in a legally required disclaimer section. RuntimeWire notes Cerebras completed a $6.4 billion gross IPO in May 2026, so roadmap credibility now carries public-market consequences, raising the incentive to lead with multiples and dated targets. Named prefill partners and customer references also serve commercial positioning.
Consistent but narrowly sourced
Two independent publishers agree on all overlapping specifications with no contradictions, and one systematically labels the provenance of each figure, which supports confidence in what was announced. Confidence in performance and operational claims is lower because both accounts ultimately depend on the same vendor primary sources, only one third-party analyst view and one preview observation are available, and no deployment, pricing or field data exists to test the rack-cadence thesis.
product
Cerebras's CS-4 is three old wafers in a new rack: price the packaging, not the silicon2 distinct publishers
invest
Nvidia's $20bn Groq buy becomes shipping racks, and the tape reads it as execution risk2 distinct publishers
invest
Speed becomes a SKU: OpenAI and Google put a separate price on latency3 distinct publishers
build
Solar Pro 4 turns model routing into a procurement decision, not a research one1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 25, 2026
1 article · August 24, 2026