Build9 publishers3 min readPublished Updated
OpenAI's first Jalapeno numbers buy it leverage, not a procurement input
The 1.7x to 3.6x latency range is set by the baseline systems, not the chip, and the report's own publication date is unsettled. Read it as direction, not evidence.
The Engineer · Build desk
What happened
- OpenAI published its first measured results from working Jalapeno silicon on Tuesday.
- It reported end-to-end latency 1.7 to 3.6 times lower than the Nvidia GB200 and GB300 systems it compared against, plus 1.5 to 1.9 times more work per watt at peak throughput.
- OpenAI ran the tests itself and no one has reproduced them in full; SemiAnalysis says it watched the InferenceX runs without running the whole suite.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint Because the range endpoints are driven by how the baseline systems performed rather than by Jalapeno's own spread, no single ratio in the set maps onto a buyer's own model mix or serving pattern.
- exposure A page whose publication date is unsettled and unversioned is a weak artefact to cite in a capacity plan, and the cost of that lands on whoever quoted it, not on OpenAI.
- decision Anyone pricing multi-year inference commitments has to decide how much weight to give hardware that will not carry production traffic until the end of 2026.
- precedent Publishing on a public harness that only the vendor has executed normalizes treating an open benchmark as an independent one, which makes the next set of first-party numbers easier to wave through.
The range is not a property of the silicon. Work back through the per-model tables: GPT-OSS 120B finished in 1.03 seconds on Jalapeno against 1.80 on a GB200 [8], which is 1.75x [1]. DeepSeek R1 went from 5.99 seconds to 1.65 against a GB300 [9], or 3.63x [2]. Kimi K2.5 went 5.31 to 1.56 [10], 3.40x [3]. Jalapeno's three absolute latencies sit within 1.6x of each other while the comparison systems' span 3.3x [4]. So the top of the headline range is mostly a statement about how a GB300 behaved in OpenAI's DeepSeek R1 configuration at that operating point, and the configuration is precisely what no outside party has run [14].
One methodological choice runs against OpenAI's interest. The per-watt results were normalized on published package ratings, 700 watts for Jalapeno against 1,200 and 1,400 for the Nvidia systems [11], even though measured sustained draw stayed at or below 550 watts [12]. That rating sits 27 percent above the number the chip actually pulled [5]. Nothing in the material says what the Nvidia systems drew, so the discount is not necessarily symmetric, but on its own side of the ledger OpenAI gave away efficiency it says it had.
The 104.3x figure is a different kind of number, and worth separating from the rest [6]. It applies to one DeepSeek R1 operating point, measured as the throughput Jalapeno could hold while matching the fastest time between tokens the comparison system ever reached [7]. On the same model at peak throughput, the advantage was 1.7x [7]. That is a 61-fold gap between the biggest slide number and the ordinary one [6]. The narrow figure is not meaningless, because agent workloads run many steps in sequence and per-token delay compounds across a task [21]. It is simply not a fleet-level efficiency claim.
Then there is provenance. The live engineering report carries an August 25th, 2026 date, the supplied page metadata dates the same page to August 3rd, no public revision history establishes which came first, and OpenAI's X announcement is timestamped August 25th [15]. For a marketing page this is trivia. For a document a buyer wants to cite in a capacity plan, a moving date with no changelog means there is no fixed record of what was claimed when.
What the numbers are actually for is visible in who benefits from them existing. A production deployment gives OpenAI direct control over part of its own inference bill and a source of capacity outside its chip suppliers [18], which is the strategy Greg Brockman described in June as making compute more abundant [17]. Broadcom did silicon implementation and networking, Celestica the boards, racks and system integration [19], and the program reached tape-out nine months from initial design [20]. Richard Ho's hardware team now has to carry laboratory results into production traffic [16], with deployment planned for the end of 2026 [13] and scheduling, utilization and reliability conditions that a controlled comparison cannot settle [23].
SemiAnalysis says it watched the InferenceX runs in person but did not run the full suite itself [14]. Witnessed is not reproduced. Until someone else executes the harness on the same stack, the honest use of 1.7x to 3.6x is as a directional read on where OpenAI thinks its cost curve is going.
What to watch
- A third-party InferenceX run on the same software stack and comparison systems, published in full rather than witnessed.
- Any correction or changelog on the results page that fixes whether it first went up on August 3rd or August 25th.
- Nvidia or another accelerator vendor publishing its own InferenceX numbers on the same three models, which would give the ratios a second reference point.