Build2 publishersReports disagree2 min readPublished
OpenAI sells the same GPT-6.1 Sol model faster for six times the token price
OpenAI's Ultrafast tier bills GPT-6.1 Sol at six times its standard per-token rate to return the same model's output faster. Teams now have to pick, request by request, which agent and interactive calls are worth that much for a shorter wait.
The Engineer · Build desk

What happened
- OpenAI added GPT-6.1 Sol to the Ultrafast tier on October 8 and said access is rolling out across the API, Codex and ChatGPT Work.
- Ultrafast lists Sol at $12 per million input tokens and $60 per million output tokens, against standard API rates of $2 and $10.
- OpenAI's DevDay recap claims up to eight times faster token generation in Codex and up to six times in the API.
- Vercel's AI Gateway now accepts the Ultrafast tier for GPT 6 Astra and GPT 6.1 Sol through the AI SDK, Chat Completions and Responses APIs.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost The premium is $50 per million output tokens against $10 per million input, so agents that write long plans or code carry most of the extra bill.
- decision Standard stays the default, so engineers have to opt in each latency-sensitive request themselves, and the tier becomes a per-call choice made in application code.
- contradiction Runtimewire reports OpenAI documentation confirming EU residency for Ultrafast, while Vercel says EU-pinned requests run at standard, so EU-bound teams cannot count on the faster tier.
- constraint Build accounts get 1 million Ultrafast tokens a minute, in a pool separate from Standard and Fast, so paying for speed does not buy any extra throughput.
In the API, the price multiple and OpenAI's best-case speedup are the same number [20]. At that ceiling, a dollar buys as much generation speed on Ultrafast as on standard. Below it, the dollar buys less. Runtimewire notes that OpenAI's speed figures are stated maximums, not guaranteed response times for every request [10].
Generation is only part of what a user waits for. OpenAI recommends persistent WebSocket connections for agents that make successive tool calls, and warns that network overhead can eat into the latency benefit [6]. Vercel's changelog passes on the same advice for the Responses API [7]. An agent turn that spends most of its time on round trips and tool execution still pays six times the standard rate on every token it generates [3].
A block of one million tokens each way costs $12 on standard and $72 on Ultrafast, before other charges [8]. That is a $60 difference [21]. OpenAI introduced Sol on September 29 as a lower-cost model for complex coding and professional work [19]. Prompts over 272,000 input tokens move to different rates, and regional processing adds 10% where it is offered [12].
Opting in is easy to build. On Vercel's gateway it is one provider option, `serviceTier: 'ultrafast'` under `providerOptions.openai` [15]. Deciding which requests get it is harder. OpenAI's documentation recommends Ultrafast when speed justifies the higher cost [18], which is advice any premium product could give. The developer thread is more specific. It names outage debugging, agents navigating applications and live experiences [4]. According to Runtimewire, those are intended use cases, not independently measured customer results [5]. I think the tier fits calls where a person is watching tokens arrive. Background agent work belongs on standard until a test on your own traffic shows otherwise.
The fallback rule is good engineering. On Vercel's gateway, a request that falls back to another tier is billed at the rate of the tier that actually served it, so a request that drops to standard pays standard prices [17]. Because price and speed both follow the tier that served the request, a latency comparison has to log which tier answered each one. If it does not, the fast sample will include standard responses.
Access is narrower outside the API. In Codex and ChatGPT Work, Ultrafast is limited to Pro 500, eligible usage-based Enterprise and credit-based Edu plans, and Enterprise administrators must enable it [11].
What to watch
- Whether OpenAI and Vercel reconcile EU processing for Ultrafast; a confirmed EU path on the gateway would open the tier to residency-bound teams.
- Independent latency measurements of Ultrafast API requests against OpenAI's up-to-six-times ceiling, including agent turns with tool calls.
- Whether OpenAI keeps the six-times multiple as Ultrafast extends beyond GPT 6 Astra and GPT 6.1 Sol.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence58
- Adoption20
- Hype gap+15
- Incentives65
- Confidence55
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
OpenAI added GPT-6.1 Sol to its Ultrafast service tier on October 8; its developer account said access is rolling out across the API, Codex and ChatGPT Work.
- [2]
GPT-6.1 Sol Ultrafast costs $12 per million input tokens and $60 per million output tokens in the API; Sol's standard rates are $2 per million input tokens and $10 per million output tokens.
- [3]
Requests served at Ultrafast are billed at 6x the standard per-token rate.
- [4]
OpenAI's October 8 thread names outage debugging, agents navigating applications and live experiences as uses where latency has consequences.
- [5]
The thread's examples describe intended use cases, not independently measured customer results.
- [6]
OpenAI recommends persistent WebSocket connections for agentic applications that make successive tool calls, warning that network overhead can eat into the latency benefit.
- [7]
For workflows with frequent tool calls, OpenAI recommends the Responses API over a persistent WebSocket connection to reduce overhead between turns.
- [8]
At standard rates, one million input tokens plus one million output tokens costs $12; the same volume costs $72 in Ultrafast, before any other charges.
- [9]
OpenAI's DevDay recap specifies up to eight times faster token generation in Codex and up to six times in the API; the October 8 thread described speeds up to eight times faster than Sol Standard.
- [10]
The speed figures are OpenAI's stated maximums, not guaranteed response times for every request.
- [11]
OpenAI says API access to Ultrafast is available to all users, while access in Codex and ChatGPT Work is limited to Pro 500, eligible usage-based Enterprise and credit-based Edu plans; Enterprise administrators must enable it.
- [12]
The model page says prompts exceeding 272,000 input tokens use different rates, and regional processing adds a 10% premium where available.
- [13]
OpenAI's API guide lists GPT-6.1 Sol Ultrafast rate limits of 1 million tokens per minute for Build, 4 million for Launch and 40 million for Grow, separate from Standard and Fast limits.
- [14]
Vercel's AI Gateway supports OpenAI's Ultrafast service tier for GPT 6 Astra and GPT 6.1 Sol, requested through AI SDK, the Chat Completions API or the Responses API.
- [15]
Vercel's example enables the tier with providerOptions: { openai: { serviceTier: 'ultrafast' } }.
- [16]
Standard processing remains the default when no service tier is specified.
- [17]
Requests that fall back to another tier are billed at the rate for the tier actually served.
- [18]
OpenAI's Ultrafast documentation calls it the API's fastest service tier and recommends it when speed justifies the higher cost.
- [19]
OpenAI introduced GPT-6.1 Sol on September 29 and positioned it as a lower-cost model for complex coding and professional work.
- [20]
In the API, Ultrafast's 6x price multiple equals OpenAI's stated maximum 6x generation speedup, so at best generation speed per dollar matches standard, and below the ceiling it is lower.
- [21]
Ultrafast costs $60 more than standard for one million input plus one million output tokens.
- [22]
The Ultrafast premium over standard is $50 per million output tokens and $10 per million input tokens.
- [23]
According to Runtimewire, OpenAI's documentation confirms GPT-6.1 Sol Ultrafast supports US and EU data residency, alongside global processing.
ReportedContestedSource: Runtimewire, citing OpenAI documentation2 sources— create a free account to open themView cited source - [24]
On AI Gateway, Ultrafast supports US and global processing; requests pinned to unsupported regions, such as the EU, run at the standard (default) tier.
Sources
2 independent publishers whose own reporting we read for this story.
- blog.vercel.comOpenAI Ultrafast mode now available on AI Gateway
1 article · October 7, 2026
- runtimewire.comOpenAI adds GPT-6.1 Sol Ultrafast at six times standard API prices
1 article · October 8, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.