Skip to content

Build2 publishers3 min readPublished

OpenAI opens its 28-day Codex pledge by speeding GPT-6 Astra and Sol to about 50 tokens a second

OpenAI's Thibault Sottiaux says GPT-6 Astra and GPT-6.1 Sol now stream about 50 tokens a second, up from 30, as day one of a 28-day ship-or-reset pledge. The figure is an approximate streaming rate, so agent run times improve only as far as text generation dominates each loop.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying OpenAI opens its 28-day Codex pledge by speeding GPT-6 Astra and Sol to about 50 tokens a second
Generated illustration

What happened

  • Subscribers using Sign in with ChatGPT get the faster models inside third-party agents including OpenCode, Pi, Amp and Devin.
  • Two days before the pledge, Sottiaux apologized for GPT-6.1 Sol's slow start, blamed a load spike, and reset usage for paid ChatGPT accounts.
  • From October 30, the $200 Pro plan's Codex and ChatGPT Work allowance falls from 20 times the Plus allowance to 10 times.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Every agent connected through Sign in with ChatGPT draws on one plan allowance, so faster streaming changes how quickly work arrives without enlarging how much work the plan pays for.
  • decision Teams choosing a third-party agent on a ChatGPT plan have to time whole tasks on their own code, because the published figure covers text streaming and not tool calls or test runs.
  • cost Pro subscribers keep paying $200 for half the multiplier after October 30, so heavy agent users face an upgrade decision inside the pledge window.
  • precedent By counting a default speed change as its first "clear improvement", OpenAI has set the bar the remaining 27 days must clear before a reset is owed.

Tokens per second is a streaming rate. A coding agent also plans, calls tools, runs code and checks its work, and output speed does not measure that full cycle, as RuntimeWire pointed out [10]. At the quoted rates, a fixed block of output now streams in about 60% of the time it used to [22]. That cut applies only to the generation share of each loop. A step that waits on a slow test suite gains little.

For the 50 tokens a second to transfer to a given team, the agent's loop has to be bound by generation, and the rate has to hold on that team's product, plan and task mix. Sottiaux called both rates approximate [5]. He also said the models use a highly optimized tokenizer and need fewer tokens per task [6]. OpenAI did not say how it measured the rates, whether they hold across products, plans and tasks, how many tokens the tokenizer saves, or whether it changed usage limits, the models or per-request compute [5][6][7]. The announcement undersells its own figures. The "about 50%" OpenAI put on the change [1] is smaller than the roughly 67% that 30 to 50 implies [21].

The delivery path is good engineering. Sign in with ChatGPT lets an eligible user run a participating app on their ChatGPT plan, with usage counted against the plan and no API key [11]. The speedup shipped as a default, so it reached those apps with no setting to change [4]. Sottiaux said users should feel it within two hours of his October 5 thread on X [9]. The allowance is shared the same way. Users can set weekly limits per app, but an app's limit does not create a separate pool of usage [12].

Resets work as the penalty because OpenAI has made them routine. Sottiaux reset Codex limits for all paid plans in April. He reset usage twice in one day in July after multi-agent regressions, and again when Codex and ChatGPT Work passed 8 million active users. He and Sam Altman handed out a banked reset on stage at DevDay on September 29 [17]. An unofficial tracker had a 28-day scoreboard up within hours of the pledge [18]. One developer quoted by The New Stack suggested burning the allowance every day and hoping OpenAI misses a ship [19]. Mark Kretschmann, a software engineer at ASYS Group, read the pledge as 28 Codex resets in 28 days [20]. "Let the improvements begin," Sottiaux wrote [16].

OpenAI announced the Pro cut alongside a new $500 plan. That plan offers 25 times the Plus allowance and an Ultrafast tier for GPT-6 Astra at up to eight times standard speed in Codex [15]. I'd expect that tier to matter more to latency-sensitive teams than day one did, if the eight times holds on their workloads. For a team choosing among agents on a ChatGPT plan before the window closes, I'd measure task completion time on its own repositories, and how fast each tool drains the shared allowance.

What to watch

  • Whether any day in the 28-day window misses a ship and triggers the full usage reset Sottiaux promised.
  • Independent task-completion timings for GPT-6 Astra and GPT-6.1 Sol inside OpenCode, Amp or Devin at the new default rate.
  • Whether the Ultrafast Astra tier's up-to-eight-times speed holds in Codex once the $500 plan and the October 30 Pro cut take effect.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories