Security1 distinct publisher3 min readUpdated
OpenAI's Ultrafast mode for GPT-5.6 Sol, running on Cerebras, is up to 14 times faster than standard processing. The company is already using it for its own incident response, which cuts both ways.
The Watch · Security desk

Compiled by The WatchSomething wrong?How this is made
OpenAI has put GPT-5.6 Sol into a limited preview "Ultrafast" mode for a select group of customers, launching first through its API, and says it runs up to 14 times faster than standard processing at up to 750 output tokens per second [1][2][3]. The mode is powered by Cerebras under the two companies' partnership on ultra-low-latency inference [4], and the consequence for security teams is not the multiplier itself but what it does to the shape of a work cycle.
750 output tokens per second works out to roughly 45,000 tokens a minute [13]. Divide the ceiling by the claimed speedup and standard processing implies something near 54 tokens per second [14]. Both figures are ceilings, not floors: OpenAI describes them as "up to" [2][3]. The launch material's illustration is a side-by-side build of a working 3D warehouse simulator from the same text prompt [12], which conveys wall-clock feel and says nothing about sustained throughput under contention.
The security-relevant detail is OpenAI's account of its own use. The company says it is running Ultrafast for incident response, including reading logs, analyzing traces, synthesizing conversations, identifying follow-up checks, and helping prepare or validate fixes [8]. Its stated rationale is that faster inference shortens the time between observing a signal, testing a hypothesis, and selecting the next action, with engineers still responsible for judgment and deployment [9]. That last clause is the load-bearing one. A pipeline that produces candidate hypotheses at 45,000 tokens a minute [13] does not make the human reviewer faster; it makes the reviewer the queue.
Most automation built over the past two years assumed the model was the slow part. Rate limits, batch windows, nightly enrichment jobs and "one analyst can approve everything the bot proposes" staffing all encode that assumption. OpenAI's own example of the change is research workflows involving knowledge searches, data queries, connected tools and synthesis [10], where teams that would launch experiments overnight and review results the next morning can instead complete multiple iterations during the workday [11]. Anyone who has sized a detection-tuning or triage loop around one review cycle per day should expect that budget to be wrong in the same direction.
The symmetry is unpleasant and unaddressed. The same compression from overnight-to-morning down to several passes inside a working day [11] applies to any loop that iterates against a target, and interactive latency is what makes multi-step tool use practical rather than something you kick off and check later. The announcement as published discusses value assessment and product change [6] and does not mention misuse, abuse mitigation, or rate limiting [16]. Preview customers are already testing across coding, commerce, financial research, support and other interactive applications in production environments [5], so this is not a sandbox question. John Crepezzi of Jane Street, quoted by OpenAI, said the speed increase "enables different ways of using the models" and makes it practical for developers to work alongside them more productively [7], which is a fair description of both the intended and the unintended user.
Watch three things: whether the 750 tokens per second ceiling [3] survives general availability and concurrency, whether your logging and audit telemetry can record agent actions at that rate without sampling them away, and what OpenAI publishes from the program it says will assess where an order-of-magnitude speed increase delivers the most value [6]. Any control that depends on a human reading output in real time needs re-measuring before the preview widens.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
OpenAI says the service runs up to 14 times faster than Standard processing.
OpenAI says the service generates up to 750 output tokens per second.
During the preview, customers are testing Ultrafast across coding, commerce, financial research, support, and other interactive applications in production environments.
John Crepezzi, AI Assistants, Jane Street, said: "The increase in speed brought by Cerebras is impressive. It enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them."
OpenAI is applying Ultrafast to research workflows involving knowledge searches, data queries, connected tools, and information synthesis.
OpenAI said research teams that might normally launch experiments overnight and review the results the following morning could instead complete multiple iterations during the workday.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single-source vendor relay, no independent measurement
All substance derives from one publisher summarizing OpenAI's own announcement. The headline performance figures are attributed ('up to 14x', 'up to 750 output tokens per second') with no methodology, prompt set, latency distribution, or quality comparison, and no third-party benchmark exists in the cluster. The internal-use and research-workflow benefits are described qualitatively or conditionally rather than measured, and the demonstration evidence is a vendor-supplied side-by-side illustration.
Limited preview with production testing but no disclosed scale
Real deployment signals exist: an API-first limited preview, preview customers reported to be testing in production across several workload categories, one named user (Jane Street), and OpenAI's own internal incident-response usage. But access is gated to a select group, no cohort size, volume, pricing, or GA timing is disclosed, and no usage metrics are published, so measurable adoption remains early-stage.
Order-of-magnitude framing outruns published proof
The framing - 'up to 14x', 750 tokens/sec, order-of-magnitude speed, agent work moving from overnight to interactive - is materially stronger than the evidence supplied. Both peak figures are ceilings from the vendor with no methodology, the demo is an illustration, the workflow gains are conditional, and cost, quality-at-speed, quota, and abuse-mitigation questions are simply absent. Gap is positive but not extreme, because concrete first-party and named-customer deployments do exist behind the rhetoric.
Vendor-originated narrative with aligned partner and customer voices
Every load-bearing statement originates with OpenAI, which is launching a paid API tier, and the performance is credited to Cerebras, its named inference partner - both parties benefit commercially from an order-of-magnitude speed narrative. The only external voice is a customer quote distributed inside the same announcement, and the illustration is vendor-supplied. The publisher adds no adversarial scrutiny or independent testing, so no counter-incentive is present in the cluster.
Clear provenance, narrow base
Confidence in what was said is high: the source is unambiguous about attribution, figures, and use cases, and the derived arithmetic is straightforward. Confidence in the underlying performance, durability, and scale is low because there is one publisher, one vendor origin, no benchmark, and no disclosed pricing or availability path.
build
OpenAI puts latency on the price list: 750 tokens/sec, gated by workload fit3 distinct publishers
invest
Speed becomes a SKU: OpenAI and Google put a separate price on latency3 distinct publishers
product
Cerebras's CS-4 is three old wafers in a new rack: price the packaging, not the silicon2 distinct publishers
build
Grok 4.6 lands in Copilot two days after launch, and the model picker becomes a procurement problem1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 14, 2026