Published Build3 min read
OpenAI Puts Latency on the Price List
Ultrafast runs GPT-5.6 Sol at up to 750 output tokens per second on Cerebras silicon. The interesting part is that speed is now a tier you buy, not a compromise you accept.
Written for builders.See today for builders
What happened
- OpenAI opened a limited preview of Ultrafast, a new API service tier, on August 13th.
- The Ultrafast tier runs GPT-5.6 Sol on Cerebras inference hardware, generating as many as 750 output tokens per second.
- OpenAI says Ultrafast can run GPT-5.6 Sol up to 14 times faster than its Standard processing tier.
- The rollout turns inference speed into a separate product tier rather than reserving the fastest responses for smaller models.
- A 14x speed multiple over the Standard tier implies a Standard-tier output rate of roughly 54 tokens per second.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
OpenAI opened a limited preview of Ultrafast on August 13th, an API service tier that the company says runs its flagship GPT-5.6 Sol at up to 750 output tokens per second on Cerebras hardware, up to 14 times the speed of its Standard processing tier [1][2][3]. The consequence is not the headline number: it is that response speed is now a separate product line rather than something you obtained by downgrading to a smaller model [4].
Run the 14x claim backwards and the implied Standard-tier rate is roughly 54 output tokens per second [5]. On a 1,500-token answer, that is about two seconds against about twenty-eight [6]. Twenty-six seconds is the gap between a loop a human stays inside and a loop a human abandons, and it is the reason several application shapes have been off the table. OpenAI is pointing the tier at exactly those shapes: incident response, voice agents, financial research, customer support, commerce and interactive scientific work [7]. Its own framing is that developers can redesign an application around lower latency, while a faster chat response mostly changes how an existing interface feels [8].
The named early testers are Jane Street, Podium, Basis and Rogo [9]. Podium product lead Courtland Lykins said Ultrafast had been valuable in the company's voice stack because "the speed completely changes the call experience" for complex work [10]. Rogo told OpenAI the tier made complex analysis feel closer to real-time interaction [11]. Internally, OpenAI says engineers are using it to read logs, traces and discussions during outages, and researchers are trying to compress experiment cycles that previously ran overnight into repeated same-day iterations [12][13].
Now the caveats, which matter more than the testimonials. The 750-token figure measures output generation and does not by itself establish time to first token or the wall-clock cost of a workflow that calls tools and external systems [14]. The 14x comparison is OpenAI's own measurement against its own Standard tier [15]. Real production numbers will move with prompts, reasoning settings, output length and the surrounding application [16]. And the schedule has already stretched: OpenAI disclosed the 750-token service in its June 26th GPT-5.6 Sol preview and said it would launch in July, the model reached general availability on July 9th, and the Cerebras-backed tier entered limited preview on August 13th, 35 days later [17][18][19]. Access remains restricted to selected customers, widening as compute capacity grows [20].
The supply side is a January 14th agreement to add 750 megawatts of low-latency capacity in stages through 2028, from a December 2025 master relationship agreement that also gave OpenAI an option on another 1.25 gigawatts by the end of 2030 [21][22]. Cerebras attributes the throughput to its wafer-scale design, which keeps model weights in on-chip memory [23][24].
OpenAI is also on both sides of the trade. In July it exercised warrants for 10,033,508 nonvoting Class N shares at $0.00001 each, a cash cost of about $100.34, for 4.22% of Cerebras [25][26][27]. Cerebras disclosed it on page 29 of a quarterly filing after the close on August 12th; at the roughly $229 Class A price the stake implied about $2.3 billion [28][29]. Warrants for 23,411,518 more shares remain, taking OpenAI to about 12.8% if all vest [30][31]. Cerebras shares fell on August 13th anyway [32].
Watch for a published time-to-first-token figure, for general availability rather than allowlist, and for what the tier costs relative to Standard. Until then, treat Ultrafast as a capacity experiment you can prototype against, not a latency budget you can promise a customer.
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
OpenAI opened a limited preview of Ultrafast, a new API service tier, on August 13th.
ReportedView cited source - [2]
The Ultrafast tier runs GPT-5.6 Sol on Cerebras inference hardware, generating as many as 750 output tokens per second.
ReportedView cited source - [3]
OpenAI says Ultrafast can run GPT-5.6 Sol up to 14 times faster than its Standard processing tier.
- [4]
The rollout turns inference speed into a separate product tier rather than reserving the fastest responses for smaller models.
ReportedView cited source - [7]
OpenAI is positioning Ultrafast for applications where users or automated systems are waiting on an answer: incident response, voice agents, financial research, customer support, commerce and interactive scientific work.
- [8]
OpenAI is releasing the tier through the API first on the reasoning that developers can redesign an application around lower latency, while a faster chat response mainly changes how an existing interface feels; cited examples include voice agents sustaining conversation with fewer pauses, coding tools returning larger edits, and incident-response systems processing changing evidence before an outage moves into its next phase.
Sources & coverage · 1 publisher
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- runtimewire.comRuntimeWire StaffAug 13OpenAI previews GPT-5.6 Sol at up to 14x standard speed
- runtimewire.comRuntimeWire StaffAug 13OpenAI acquired 4.2% of Cerebras before GPT-5.6 Ultrafast launch
Additional citations
- OpenAI
- Courtland Lykins, Podium product lead
- Rogo, via OpenAI
- Cerebras

