Security1 publisher3 min readPublished
750 tokens a second moves agents from overnight batch to interactive latency
OpenAI's Ultrafast mode for GPT-5.6 Sol, running on Cerebras, is up to 14 times faster than standard processing. The company is already using it for its own incident response, which cuts both ways.
The Watch · Security desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- OpenAI's GPT-5.6 Sol on Ultrafast mode is available in limited preview to a select group of customers, launching first through the OpenAI API.
- OpenAI says the service runs up to 14 times faster than Standard processing.
- OpenAI says the service generates up to 750 output tokens per second.
- Ultrafast is powered by Cerebras as part of the companies' partnership on ultra-low-latency inference.
- During the preview, customers are testing Ultrafast across coding, commerce, financial research, support, and other interactive applications in production environments.
Compiled by The WatchSomething wrong?How this is made
Why it matters
OpenAI has put GPT-5.6 Sol into a limited preview "Ultrafast" mode for a select group of customers, launching first through its API, and says it runs up to 14 times faster than standard processing at up to 750 output tokens per second [1][2][3]. The mode is powered by Cerebras under the two companies' partnership on ultra-low-latency inference [4], and the consequence for security teams is not the multiplier itself but what it does to the shape of a work cycle.
750 output tokens per second works out to roughly 45,000 tokens a minute [13]. Divide the ceiling by the claimed speedup and standard processing implies something near 54 tokens per second [14]. Both figures are ceilings, not floors: OpenAI describes them as "up to" [2][3]. The launch material's illustration is a side-by-side build of a working 3D warehouse simulator from the same text prompt [12], which conveys wall-clock feel and says nothing about sustained throughput under contention.
The security-relevant detail is OpenAI's account of its own use. The company says it is running Ultrafast for incident response, including reading logs, analyzing traces, synthesizing conversations, identifying follow-up checks, and helping prepare or validate fixes [8]. Its stated rationale is that faster inference shortens the time between observing a signal, testing a hypothesis, and selecting the next action, with engineers still responsible for judgment and deployment [9]. That last clause is the load-bearing one. A pipeline that produces candidate hypotheses at 45,000 tokens a minute [13] does not make the human reviewer faster; it makes the reviewer the queue.
Most automation built over the past two years assumed the model was the slow part. Rate limits, batch windows, nightly enrichment jobs and "one analyst can approve everything the bot proposes" staffing all encode that assumption. OpenAI's own example of the change is research workflows involving knowledge searches, data queries, connected tools and synthesis [10], where teams that would launch experiments overnight and review results the next morning can instead complete multiple iterations during the workday [11]. Anyone who has sized a detection-tuning or triage loop around one review cycle per day should expect that budget to be wrong in the same direction.
The symmetry is unpleasant and unaddressed. The same compression from overnight-to-morning down to several passes inside a working day [11] applies to any loop that iterates against a target, and interactive latency is what makes multi-step tool use practical rather than something you kick off and check later. The announcement as published discusses value assessment and product change [6] and does not mention misuse, abuse mitigation, or rate limiting [16]. Preview customers are already testing across coding, commerce, financial research, support and other interactive applications in production environments [5], so this is not a sandbox question. John Crepezzi of Jane Street, quoted by OpenAI, said the speed increase "enables different ways of using the models" and makes it practical for developers to work alongside them more productively [7], which is a fair description of both the intended and the unintended user.
Watch three things: whether the 750 tokens per second ceiling [3] survives general availability and concurrency, whether your logging and audit telemetry can record agent actions at that rate without sampling them away, and what OpenAI publishes from the program it says will assess where an order-of-magnitude speed increase delivers the most value [6]. Any control that depends on a human reading output in real time needs re-measuring before the preview widens.