Build1 publisher2 min readPublished
OpenAI books 750MW of Cerebras inference capacity in three 250MW segments through 2028
OpenAI has contracted 750MW of Cerebras wafer-scale inference capacity, delivered in three 250MW segments by the end of 2028. The dates are contractual, but neither company has published pricing, latency figures or a plan for which customers get access.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The terms sit in a Master Relationship Agreement effective December 24, 2025, according to regulatory filings summarised in a dev.to post.
- The agreement also allows for extra capacity, a hardware purchase path, service-level terms and contemplated exclusivity between the two companies.
- An SEC filing shows OpenAI introduced a Codex Spark model running on Cerebras infrastructure around February 12, 2026.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint Near-term capacity is the first 250MW segment, a third of the headline figure, so a 2026 plan sized against the full 750MW is two segments early.
- decision A team that wants Cerebras-backed latency now has one documented model to target, Codex Spark, and cannot yet pick a tier or product known to run on the hardware.
- exposure The service-level terms bind Cerebras to OpenAI, so OpenAI customers get no stated response-time or availability commitment from the deal and still carry that risk in their own designs.
Nobody runs a prompt in megawatts. The contract counts power [4]. Power sets how much Cerebras hardware OpenAI can switch on. It does not set tokens per second for a given model. Cerebras described the deployment as meant to serve OpenAI customers and support faster, real-time interactions [2]. Neither company has published model-by-model performance figures, pricing or a detailed customer access plan [10].
The contract detail is secondhand. It comes from regulatory filings as summarised in a dev.to post that closes with a pitch for Scalevise's AI consultancy [13]. On that account the structure is sound, because staging 750MW in dated blocks lets one block slip without stalling the others. Each of the three 250MW segments carries an end-of-year milestone, and the first delivery target falls in 2026 [5].
For 2026 planning, the figure to use is 250MW. The first segment is a third of the total [1]. Full capacity arrives about three years after the agreement took effect on December 24, 2025 [2]. OpenAI's public wording is looser: it describes capacity coming online in multiple tranches through 2028 [7].
The agreement covers more than delivery dates. It leaves room for more capacity, sets out a route to buying hardware, and includes service-level terms and exclusivity that the two companies have contemplated [6]. Those service-level terms sit between OpenAI and Cerebras. The post warns against reading the capacity as a promise of a specific response time, price reduction or availability level for every OpenAI customer [12].
So far the operational evidence is one model. A separate SEC filing shows an OpenAI Codex Spark model powered by Cerebras infrastructure, introduced around February 12, 2026 [8]. According to the post, that filing shows early use and nothing broader. It does not establish that every OpenAI model or service runs on Cerebras capacity [9].
That limits how far any speed claim transfers. For a team's response times to improve, its traffic has to land on a model OpenAI serves from this hardware. The one documented case today is Codex Spark [8].
What to watch
- OpenAI naming which models, products or customer tiers route to Cerebras capacity, and at what price.
- Whether the first 250MW segment meets its 2026 year-end milestone.
- Filing detail on the contemplated exclusivity, to see whether it limits Cerebras's other customers or OpenAI's other suppliers.