Published · 4h agoBuild3 min read
OpenAI's first Jalapeno numbers buy it leverage, not a procurement input
The 1.7x to 3.6x latency range is set by the baseline systems, not the chip, and the report's own publication date is unsettled. Read it as direction, not evidence.
Written for builders.See today for builders
What happened
- OpenAI published its first measured results from working Jalapeno silicon on Tuesday.
- It reported end-to-end latency 1.7 to 3.6 times lower than the Nvidia GB200 and GB300 systems it compared against, plus 1.5 to 1.9 times more work per watt at peak throughput.
- OpenAI ran the tests itself and no one has reproduced them in full; SemiAnalysis says it watched the InferenceX runs without running the whole suite.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraintBecause the range endpoints are driven by how the baseline systems performed rather than by Jalapeno's own spread, no single ratio in the set maps onto a buyer's own model mix or serving pattern.
- exposureA page whose publication date is unsettled and unversioned is a weak artefact to cite in a capacity plan, and the cost of that lands on whoever quoted it, not on OpenAI.
- decisionAnyone pricing multi-year inference commitments has to decide how much weight to give hardware that will not carry production traffic until the end of 2026.
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
OpenAI published the first measured results for its Jalapeno custom inference chip on Tuesday.
ReportedView cited source - [2]
The tests covered GPT-OSS 120B, DeepSeek R1 670B and Moonshot AI's Kimi K2.5 1T using InferenceX, SemiAnalysis' open-source inference benchmark, against commercially available Nvidia systems.
ReportedView cited source - [3]
OpenAI reported end-to-end latency between 1.7 and 3.6 times lower than the Nvidia comparison systems.
ReportedView cited source - [4]
OpenAI reported 1.5 to 1.9 times more AI work per watt at peak throughput than Nvidia GB200 and GB300 systems across the three models.
ReportedView cited source - [5]
OpenAI reported 2.1 to 4.1 times higher interactivity across the tested models.
ReportedView cited source - [6]
OpenAI reported 8.6x to 104.3x higher performance per watt at a prior-best time between tokens.
ReportedView cited source
Sources & coverage · 4 publishers
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- thenewstack.ioAmanda Caswell11h agoOpenAI built a chip in nine months. Then it let AI rewrite the code.
- runtimewire.comRyan Merket7h agoOpenAI says Jalapeno beats Nvidia Blackwell on inference speed per watt


