Build1 distinct publisher3 min readPublished
OpenAI published a benchmark win on three large models, but the figure a GPU vendor has to price is the 10-gigawatt Broadcom program running through 2029, with racks targeted from the second half of 2026.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
A purchase-order lever does not need the benchmark to survive contact with production. It needs the internal alternative to be dated and specific. The Broadcom program supplies both: 10 gigawatts of custom accelerators announced in October 2025, racks targeted to start deploying in the second half of 2026, the program running through 2029 [3]. The dev.to analysis is careful to note that those gigawatts are a roadmap and not deployed capacity [4], though a roadmap on this schedule is still enough to price a renewal. From the October 2025 announcement to the end-of-2026 initial deployment OpenAI describes is about fourteen months [19], which sits inside the term of any large GPU commitment being negotiated now.
OpenAI calls the part its first "Intelligence Processor"; the plainer description in the same piece is a custom ASIC for LLM inference, co-developed with Broadcom and turned into boards, racks and production systems with Celestica [1][2]. The narrower job description is the whole argument. A merchant GPU trains and serves; this design targets serving, with attention to memory movement, network delay, and the changing mix of prompt processing and token generation [21].
The mechanism worth reading closely is the pooling decision. Prefill is compute-heavy. Decode waits on memory bandwidth, because the system re-reads weights and the KV cache on every token [10]. Split your fleet into a prefill pool and a decode pool and the ratio is only correct while traffic matches the plan, and prompt length, output length, cache-hit rate, concurrency and speculative-decoding acceptance all move through the day, so a fixed split leaves one pool idle while the other queues [15]. OpenAI runs one flexible accelerator pool instead [14], accepts some local inefficiency to keep every device available for the next request, and avoids shipping a growing KV cache across the network between phases [16]. Cores and HBM are organised into local slices so software can place weights and KV state near the compute that needs them [11], a dedicated collective network carries the known high-bandwidth traffic between slices [12], and support for smaller matrix shapes avoids the padding and utilisation cliffs of large systolic arrays [13]. That is co-design in the literal sense, and it is good work.
It also tells you what the benchmark is a claim about. OpenAI reports strong latency and performance per watt on three large models [7]. For that to transfer to someone else's fleet, that fleet would need OpenAI's model set, OpenAI's prompt and output length distribution, and OpenAI's cache-hit and batching behaviour, because the pooling and slicing choices are tuned to exactly that traffic. The same source lists what is not yet demonstrated: production-scale economics, long-context agent performance, and fleet reliability [8]. On the published benchmark Jalapeno beats NVIDIA, and the dev.to author declines to call it a better chip than the Blackwell platform on that evidence [6].
Scale is the other thing to check before borrowing the numbers. SemiAnalysis, as cited in the piece, describes a reticle-sized compute die, HBM4 at 15.4 TB/s, and a system connecting 2,048 accelerators across 16 racks [17]. That is 128 accelerators per rack [18]. If your serving domain does not need a 2,048-way collective, most of this architecture is answering a question you do not have.
The procurement read is the durable one. Per the dev.to analysis, the chip gives OpenAI a credible way to move repeated, high-volume inference onto hardware it controls, which changes how it buys GPUs, how much pricing power NVIDIA keeps, and how expensive it is to leave CUDA [9]. Over time the same piece expects a profitable slice of inference to move off merchant GPUs, weakening one part of the software moat [20]. That weakens the moat for OpenAI, which owns both the model and the serving stack, though everyone else still pays the porting bill.
Ranked by verification strength, evidence, and original report placement.
On the benchmark OpenAI published, Jalapeno beats NVIDIA, but the dev.to author says the evidence does not yet support the claim that it is a better chip than NVIDIA's Blackwell platform.
Early results show excellent latency and performance per watt on three large models.
Jalapeno has not yet proved production-scale economics, long-context agent performance, or fleet reliability.
Jalapeno is an inference ASIC co-developed by OpenAI and Broadcom; OpenAI calls it its first "Intelligence Processor", which the dev.to piece describes plainly as a custom ASIC for large-language-model inference.
The chip is turned into boards, racks and production systems with Celestica.
OpenAI and Broadcom announced a 10-gigawatt custom-accelerator program in October 2025, with racks targeted to start deploying in the second half of 2026 and the program running through 2029.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 30, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
product
The $500bn compute asset class rests on a depreciation curve Nvidia once denied1 distinct publisher
build
A benchmark that replays real agent sessions gives back less of the generational win2 distinct publishers
build
SemiAnalysis to software teams: your token cost starts at the fab, not the price list1 distinct publisher
product
Swapping out the GPU leaves four more rack lines on Nvidia's invoice1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Specific numbers, all borrowed
The hardest figures in this story — 15.4 TB/s of HBM4, a reticle-sized die, 2,048 accelerators across 16 racks, 1.5–1.9x work per watt — reach a reader through dev.to relaying OpenAI's launch material and SemiAnalysis' architectural write-up. To its credit the piece marks the seam: the analyst firm watched the runs, OpenAI produced the numbers, the full InferenceX suite was never completed and AgentX was never attempted. Detailed enough to be falsified later, unverified today.
Samples in a lab, racks on a calendar
Nothing is in production. What exists is silicon reported to run at target frequency and power, three open models benchmarked indoors, and a schedule: racks from the second half of 2026, the program to 2029. dev.to is explicit that the gigawatts describe intent, not installed capacity — so the honest adoption reading is a design win with a delivery date, and dev.to itself lists production economics, long-context agents and fleet reliability as still open.
The piece argues down its own headline
dev.to works hard to be smaller than the story it could have written: it rejects the 'NVIDIA killer' frame outright and says the evidence does not yet make this a better chip than Blackwell. The overshoot that survives is inherited rather than manufactured — a win on an 8,000-in, 1,000-out workload and a gigawatt roadmap are being asked to stand in for a deployed fleet, and 'inference margins on the clock' arrives roughly fourteen months before the first rack is scheduled to power on.
Everyone holding a number holds a position
OpenAI published the benchmark OpenAI wins, run on a suite belonging to the analyst firm that stood in the room, concerning a chip OpenAI needs to look credible while it negotiates its next round of GPU purchases. dev.to discloses the chain instead of laundering it — 'SemiAnalysis observed the runs in OpenAI's lab, but OpenAI supplied the numbers' — which is the correct disclosure and still not an independent measurement. NVIDIA, whose margins are the subject, is not quoted.
Solid on what was built, thin on what it does to anyone
We can say with reasonable assurance what OpenAI has designed, why the serving architecture is shaped that way, and what schedule OpenAI is committing to in public. We cannot say what it costs, what NVIDIA will concede, or whether the part survives contact with long agent sessions. One publisher, no rebuttal, no unit economics, and a deployment date distant enough for any of it to slip.