Published · 5d agoScience2 min read
750 tokens per second: what Cerebras's 14x claim for GPT-5.6 Sol actually buys
The peak output rate on OpenAI's new Ultrafast tier comes from holding weights in on-chip SRAM. End-to-end gains on benchmark work land nearer 6x, and "same intelligence" is a benchmark claim.
Written for builders.See today for builders

What happened
- Cerebras announced on Aug. 13, 2026 that it is powering Ultrafast mode, a new service tier in the OpenAI API for GPT-5.6 Sol, running at up to 750 output tokens per second and up to 14x faster than Standard processing.
- Ultrafast is available initially in limited preview to OpenAI customers.
- Cerebras says Ultrafast's speed comes from its Wafer-Scale Engine architecture, which keeps model weights on-chip in 44 GB of SRAM on each wafer-sized chip rather than shuttling them between on-chip memory and off-chip storage as GPU-based inference must, eliminating the memory-bandwidth bottleneck that constrains frontier-model inference speed on conventional hardware.
- On GDP-Val, a benchmark of economically valuable knowledge-work tasks such as legal briefs, financial models and engineering reports, Cerebras reports Ultrafast delivered a 5.6x end-to-end speedup with no loss in quality.
- The reported 5.6x end-to-end GDP-Val speedup is about 40 percent of the headline 14x figure.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
Cerebras said on Aug 13, 2026 that it is powering Ultrafast, a new service tier in the OpenAI API for GPT-5.6 Sol, running the model at up to 750 output tokens per second and up to 14x faster than Standard processing [1]. The tier is in limited preview to OpenAI customers [2]. The figure is a peak output token rate, and the gap between it and measured task times is where the decision sits.
Cerebras attributes the speed to its Wafer-Scale Engine architecture, which keeps model weights on-chip in 44 GB of SRAM per wafer-sized chip rather than shuttling them to off-chip storage as GPU-based inference must, removing the memory-bandwidth bottleneck [3]. That is a ceiling claim, not a workload claim. On GDP-Val, a benchmark of knowledge-work tasks such as legal briefs and financial models, the company reports a 5.6x end-to-end speedup with no loss in quality [4] -- about 40 percent of the headline multiplier [5]. On Humanity's Last Exam, 2,500 questions, Ultrafast finished the set in just over 11 hours against more than three days for Claude Fable 5, at comparable accuracy nearly 7x faster [6]. Taken at face value, 750 tokens per second at 14x implies a Standard baseline near 54 tokens per second [7]. The Anthropic comparisons rest on output speeds reported by Artificial Analysis [8].
The load-bearing marketing claim is that Ultrafast runs with the same intelligence as Standard [9]. Same weights is not the same outputs. Work published on arXiv finds that changing evaluation batch size, GPU count or GPU version can produce significant differences in generated responses, most acutely in reasoning models [10], with one distilled 7B reasoning model showing up to 9 percent accuracy variation and 9,000 tokens of response-length difference [11]; the root cause is non-associative floating-point arithmetic at limited precision [12]. Cerebras is not a GPU, so those specific numbers do not transfer, but arithmetic order changes whenever hardware does.
What would move the number: OpenAI's Sachin Katti says the company is starting with a small group of customers to learn where the speed creates meaningful value, and will use that to inform expansion [13]. Watch for per-task accuracy deltas against Standard rather than "comparable," and for whether the 5.6x end-to-end figure holds as capacity opens.
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Cerebras announced on Aug. 13, 2026 that it is powering Ultrafast mode, a new service tier in the OpenAI API for GPT-5.6 Sol, running at up to 750 output tokens per second and up to 14x faster than Standard processing.
ReportedView cited source - [2]
Ultrafast is available initially in limited preview to OpenAI customers.
ReportedView cited source - [3]
Cerebras says Ultrafast's speed comes from its Wafer-Scale Engine architecture, which keeps model weights on-chip in 44 GB of SRAM on each wafer-sized chip rather than shuttling them between on-chip memory and off-chip storage as GPU-based inference must, eliminating the memory-bandwidth bottleneck that constrains frontier-model inference speed on conventional hardware.
ReportedView cited source - [4]
On GDP-Val, a benchmark of economically valuable knowledge-work tasks such as legal briefs, financial models and engineering reports, Cerebras reports Ultrafast delivered a 5.6x end-to-end speedup with no loss in quality.
ReportedView cited source - [6]
On Humanity's Last Exam, a 2,500-question benchmark spanning graduate-level chemistry, economics and literature, GPT-5.6 Sol Ultrafast answered the full question set in just over 11 hours, compared with more than three days of continuous compute for Claude Fable 5, reaching comparable accuracy nearly 7x faster.
ReportedView cited source - [8]
Cerebras states that, based on output speeds for Anthropic models reported by Artificial Analysis, Ultrafast is 5x faster than Claude Opus 4.8 in Fast mode and 11x faster than Claude Fable 5.
Sources & coverage · 2 publishers
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- globenewswire.com5d agoCerebras Powers Ultrafast Mode for OpenAI’s GPT-5.6 Sol
Additional citations
- Cerebras, citing Artificial Analysis
- Sachin Katti, OpenAI


