Product2 publishers3 min readPublished
Cerebras's CS-4 is three old wafers in a new rack: price the packaging, not the silicon
The first multi-wafer Cerebras system pairs a doubled clock with rebuilt power delivery and interconnect. The compute claims rest on WSE-3 Turbo dies that are otherwise unchanged.
The Product Desk · Product desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Cerebras unveiled the CS-4 on Tuesday at its Supernova event, its first system to put three wafer-scale processors in a single rack; it ships this quarter, is pitched as an inference machine for frontier models, and Cerebras says it runs them up to 30 times faster than GPU-based systems.
- The launch lands five days after OpenAI's Ultrafast mode, which runs GPT-5.6 Sol roughly 14 times faster on Cerebras silicon, went live.
- The CS-4 is the first hardware Cerebras has shipped since its $5.55bn Nasdaq debut in May, the largest US tech listing since Snowflake.
- Each CS-4 carries three WSE-3 Turbo wafers for a combined 750 petaflops of sparse FP16 compute, 129.6 petabytes per second of memory bandwidth, and support for models above 50 trillion parameters.
- Wafer-to-wafer latency falls to two microseconds from five, the rack uses half as many components as its predecessor, and power conversion has been moved, in Cerebras's own phrasing, a hundred times closer to the processors, mounted in a removable backpack at the rear of the chassis.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
Cerebras unveiled the CS-4 on Tuesday at its Supernova event, the first time it has put three of its wafer-scale processors into a single rack, and says it will ship before the end of the quarter [1]. The processor inside is the same die at roughly twice the clock, so the honest way to evaluate this box is as density, power delivery and latency engineering rather than a new chip generation [6][7].
The rack figures are 750 petaflops of sparse FP16 compute, 129.6 petabytes per second of memory bandwidth, and support for models above 50 trillion parameters [4]. Per wafer that is 250 petaflops sparse against 25 petaflops dense, a tenfold gap between the headline and the dense number [9][3]. The WSE-3 Turbo carries the same four trillion transistors, the same 900,000 cores, the same 44GB of on-chip SRAM and the same TSMC 5nm node as the WSE-3 it replaces [6]. The Register concluded it is not new silicon but the existing die pushed from about 1.4GHz to 2.8GHz, with per-wafer compute and bandwidth both exactly doubling [7][8]. The clock ratio and the compute ratio are the same number, which is what a clock bump looks like [1].
What is new sits around the wafers. Wafer-to-wafer latency falls from five microseconds to two, a 60 percent cut [5][2]; the rack uses half as many components as its predecessor; and power conversion has moved into a removable backpack at the rear, which Cerebras describes as a hundred times closer to the processors [5]. The Register estimates 120 to 140 kilowatts per rack, roughly half what comparable AMD and Nvidia rack systems draw, which is the plausible basis for the claimed tenfold gain in throughput per watt [10].
The company's up-to-30-times-faster claim is measured as tokens per second per user on a single model, gpt-oss-120b, against unnamed GPU systems [1][9]. Chief technology officer Sean Lie put it in terms buyers can test: being 30 times faster "gives an agentic system room for more than an order of magnitude as much reasoning, verification, or tool use" [12]. Chief executive Andrew Feldman said "In AI, speed is productivity", and told Reuters the company expects to get "four times as fast between now and the end of 2027, and 20 times more throughput" [11]. This is also the first hardware since Cerebras's $5.55bn Nasdaq debut in May, and it follows OpenAI's Ultrafast mode, which runs GPT-5.6 Sol roughly 14 times faster on Cerebras silicon, by five days [2][3].
The commercial disclosure is thin. Cerebras named OpenAI, G42, MBZUAI and AWS alongside the launch but disclosed no CS-4 customer agreements and no pricing [13]. AMD's Helios rack, unveiled in July, is in the partner list, consistent with a company that says it will work with everyone in AI hardware except Nvidia [14].
Second-quarter revenue was $180.1m, up 74 percent year on year with cloud revenue nearly quadrupling, but down from $193.4m in the first quarter, a 6.9 percent sequential decline, and the margin pressure flagged in June has not lifted [15][4]. The quarter produced a GAAP net loss of $450.5m against an adjusted loss of $6.9m, a gap of $443.6m, with $25.4bn in remaining performance obligations [16][6].
Watch three things. Full-year guidance of $880m to $890m implies second-half revenue of $506.5m to $516.5m against $373.5m in the first half, about 1.36 to 1.38 times the first-half run rate, after a sequential decline [16][5]. Concentration has not gone away: G42 and MBZUAI together were around 86 percent of 2025 revenue, with the January OpenAI contract worth more than $10bn at signature and second-quarter additions including Cognition, Lovable, CrowdStrike, Block and Figma [17]. And 2027, when the next generation has to arrive on new silicon rather than a faster clock [8][18].