Skip to content

Invest1 publisher3 min readPublished

OpenAI's cost-per-task argument buys Luna room for ten failed tries before it loses on price

Sarah Friar told Goldman Sachs' technology conference that a pricier model can be cheaper when it needs fewer tries, a claim Artificial Analysis puts at an eleven-to-one price spread against seven index points running the other way.

The Investor · Invest desk

What happened

  • Reuters reported on September 9 that OpenAI CFO Sarah Friar had told Goldman Sachs' Communacopia conference two days earlier that buyers should weigh cost per completed task rather than token price.
  • OpenAI is testing contracts priced on how effective its models prove for a given company rather than on consumption, as it pushes into chip design, life sciences and finance.
  • OpenAI says its own AI took Jalapeno, a Broadcom-designed inference chip it does not sell, from conception to tapeout in nine months.
  • Cadence said on June 1 that its ChipStack AI engineer can take some RTL validation cycles from five weeks to under a day, and Google's AlphaChip is already used across Alphabet's advanced chips.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • constraint Procurement loses its one cross-vendor number, because a token price is comparable between suppliers while a completed task is defined by whoever is selling it.
  • decision Billing on effectiveness moves the cost of retries onto the seller's own accounts, so every failed attempt becomes OpenAI's margin question instead of the customer's usage bill.
  • exposure Anyone underwriting OpenAI's inference-cost advantage is underwriting a single party's measurement, since the per-watt comparison with Nvidia's parts is OpenAI's own and the test suite was watched rather than reproduced.
  • contradiction The cheaper standard task belongs to the lower-ranked model, so the argument only closes if Luna's failures are rare, and the seven-point index gap is the only outside signal on that.

Divide $2.01 by $0.18 and you have the pitch. Artificial Analysis prices a standard task on Z.ai's GLM-5.3 at $2.01 and the same task on GPT-5.6 Luna at roughly $0.18 [9], which means Luna can take eleven swings, miss ten of them, and still come in under one clean GLM-5.3 run at $1.98 against $2.01 [15]. What the division does not price is the checking. GLM-5.3 scores 45 on Artificial Analysis's Intelligence Index against Luna's 38 [10], a gap of seven points [19], so Luna is buying about 84% of the measured capability for about 9% of the money [16], and the residual sits with whoever has to notice which of the eleven attempts was correct.

Cost per completed task is a serviceable unit right up to the question of who owns the denominator. Synopsys says its generative-AI copilot cuts information-retrieval time by 40% and time-to-solution by 10 to 20 times [12]; those are cycle-time claims, not dollars per finished design, and a chip team weighing OpenAI against an incumbent toolchain has no shared unit to divide by. That is the quiet cost of retiring the token price, which never measured value well but let every vendor be compared on the same number.

Friar's July 17 framing of "Useful Intelligence per Dollar" over token price [4] and the effectiveness-based contracts OpenAI is now testing [3] run the same direction, and the Luna numbers show why. Cutting Luna's price drove close to tenfold growth in usage and took Codex to 25 million users, Friar said [8], though no figure for the size of that cut has been published, so whether the trade added revenue or bought volume with margin cannot be computed from what is on the record [18]. Outcome pricing is what a seller reaches for when the meter itself is deflating.

The chip functions as a cost lever inside OpenAI, not as a product for sale. OpenAI is not trying to displace the EDA systems engineers already run [11], and Jalapeño, designed with Broadcom, is built for internal use and not for sale [7], so the claimed 1.5 to 1.9 times AI work per watt against Nvidia's GB200 and GB300 on GPT-OSS 120B, DeepSeek R1 and Kimi K2.5 [6] lands in OpenAI's own inference bill and never in a revenue line. It also arrives as OpenAI's own measurement, with SemiAnalysis having watched the InferenceX runs without reproducing the full test suite [7].

My read is that task-based pricing is a rational answer to token deflation, and the more interesting version is that it lets OpenAI charge chip design, life sciences and finance buyers [1] far more than a per-token meter would ever justify without publishing a higher per-token price. The counter-thesis sits in the same sentence: once you bill for outcomes you own the retries and you argue with the customer over whether the task was finished, which is a consulting firm's collections problem wearing a software margin. Nine months from concept to tapeout [5] is the figure OpenAI would like underwriters holding, and the only third-party figure anywhere in the pitch is a price ratio of about eleven to one [15], which prices Luna alone; it says nothing about what a finished chip is worth.

What to watch

  • Whether the effectiveness-based contracts move from test to published terms, and which party is billed when a task is judged incomplete.
  • Whether anyone outside OpenAI reproduces the InferenceX suite behind the 1.5 to 1.9 times per-watt figure.
  • Whether Artificial Analysis or a rival benchmark starts publishing retry counts per task, the number that decides the eleven-to-one gap.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories