Published Product3 min read
OpenAI Puts Latency in Its Own Tier, and Cerebras Silicon in the Serving Stack
Ultrafast serves the same flagship model at up to 14 times the speed, which removes the intelligence-versus-latency choice that has kept agent features in demo state.
Not a builder's beat, but builders have a standing stake in it.See today for builders

What happened
- OpenAI previewed Ultrafast, a new tier of its API that runs the flagship GPT-5.6 Sol at up to 14 times the usual speed, reaching around 750 output tokens a second, on hardware built by wafer-scale chipmaker Cerebras.
- Ultrafast is not a new model but a new way to serve an existing one, leaning on Cerebras chips to strip out latency.
- OpenAI says the tier dissolves a trade-off: until now, anyone who wanted genuinely real-time responses had to drop down to a smaller, less capable model, accepting less intelligence in exchange for speed.
- An AI agent that has to think for thirty seconds before every step is a demo, whereas one that answers in the time it takes to hold a conversation starts to feel like a product.
- A rate of 750 output tokens a second at 14 times normal speed implies a baseline of roughly 54 output tokens a second.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
OpenAI has previewed Ultrafast, an API tier that runs its flagship GPT-5.6 Sol at up to 14 times the usual speed, around 750 output tokens a second, on hardware from wafer-scale chipmaker Cerebras [1]. The consequential detail is what it is not: not a smaller distilled model, but the same model served differently [2], which changes a constraint every team shipping agentic features has had to design around.
Until now, according to OpenAI's pitch as reported by TNW, anyone who wanted genuinely real-time responses had to drop down to a smaller, less capable model and accept less intelligence in exchange for speed [3]. That is not a marketing problem, it is why so much agent tooling reads as a demo. An agent that thinks for thirty seconds before every step is a demo; one that answers at conversational pace starts to behave like a product [4]. Working backwards from the stated numbers, 750 tokens a second at 14x implies a baseline of roughly 54 tokens a second [5], and on the same multiple the illustrative thirty-second step would fall to about two seconds [6].
The gating matters as much as the throughput. OpenAI opened a limited preview on 13 August to a small group of customers, with plans to widen access as capacity allows [7]. No pricing has been published [8], and TNW's assessment is that running the top model at 14 times the speed on specialist hardware is unlikely to be cheap, leaving Ultrafast a premium option for latency-obsessed cases rather than a default [9]. Treat this, for now, as a capacity story wearing a product label.
The named workloads are the ones where a wait is a business cost: incident response and debugging, financial research and fraud detection, real-time customer support and voice, and e-commerce [10]. Early testers include Jane Street, Podium, Basis and Rogo, who describe the change as qualitative rather than incremental [11], with one saying speed "completely changes the call experience for complex work" and unlocks "synchronous experiences for users that were previously limited by intelligence" [12]. These are reference customers surfaced alongside a launch, so weigh them accordingly.
The mechanism is architectural. Cerebras cuts processors the size of a dinner plate from a single silicon wafer, which lets an entire model sit on one chip instead of being split across racks of Nvidia GPUs that must constantly shuttle data between them [13]. Stripping out that internal traffic is what collapses the delay between prompt and reply, and it underpins the company's long-running claim that its design suits inference better than chips built for training [14]. For Cerebras the endorsement lands at a useful moment: it went public in one of the year's biggest listings but has since struggled to convince the market that wafer-scale ambition converts into durable profit [15]. For OpenAI, leaning on Cerebras is also a step away from total dependence on Nvidia, consistent with its own custom silicon work [16].
The wider frame is that as raw capability plateaus, the contest moves to who runs models fastest and cheapest, a shift that has lifted inference specialists such as Groq, with SambaNova and specialist clouds chasing the same prize [17].
Watch three things: whether the preview widens beyond a handful of accounts, what the price per token turns out to be, and whether Cerebras capacity, not model quality, becomes the binding constraint on real-time products.
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
OpenAI previewed Ultrafast, a new tier of its API that runs the flagship GPT-5.6 Sol at up to 14 times the usual speed, reaching around 750 output tokens a second, on hardware built by wafer-scale chipmaker Cerebras.
- [2]
Ultrafast is not a new model but a new way to serve an existing one, leaning on Cerebras chips to strip out latency.
- [3]
OpenAI says the tier dissolves a trade-off: until now, anyone who wanted genuinely real-time responses had to drop down to a smaller, less capable model, accepting less intelligence in exchange for speed.
- [4]
An AI agent that has to think for thirty seconds before every step is a demo, whereas one that answers in the time it takes to hold a conversation starts to feel like a product.
- [7]
OpenAI opened a limited preview of Ultrafast on 13 August to a small group of customers, with plans to widen access as capacity allows.
Sources & coverage · 1 publisher
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- thenextweb.comAna Maria ConstantinAug 13OpenAI’s new Ultrafast mode runs GPT-5.6 Sol 14 times faster, on Cerebras chips
Additional citations
- TNW
- OpenAI, via TNW
- unnamed early tester, via TNW



