Science1 distinct publisher3 min readUpdated
Artificial Analysis puts OpenAI's new reasoning model at three times the index score of its price-tier peers and more than twice their token appetite. The output line dominates the invoice.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
Artificial Analysis has published its run on GPT-5.6 Luna (max), the OpenAI reasoning model released on July 9, 2026 [8], scoring it 52 on the Artificial Analysis Intelligence Index against a median of 17 for reasoning models in a similar price tier [1]. The number that will actually show up on an invoice is a different one: the model generated 130M output tokens to complete the index, against a 60M median [2].
That is 2.17 times the token use of the median model in its cohort [1], for 3.06 times the score [11]. Published prices are $0.20 per 1M input tokens and $1.20 per 1M output tokens, against medians of $0.25 and $0.90 [3]. So the input side is cheap and the output side is not, and the output side is where the volume sits. At list price, 130M output tokens cost $156.00 [2]. A model with median verbosity at the median output price would have spent $54.00 on the same evaluation [3], making Luna's output bill 2.89 times larger [4]. Artificial Analysis reports the full evaluation cost at $172.17 [4], which means output tokens are roughly 90.6 percent of it [5]. The remaining $16.17 covers input and cache-hit tokens, which the benchmark does not break out [10].
The honest counterpoint is that verbosity buys something here. Per point of index score, Luna's output spend works out at $3.00 against $3.18 for the median-shaped run, about 5.5 percent cheaper [6]. If your workload is graded on getting the answer right, the extra tokens are not waste. If your workload is fixed-shape, high-volume and priced per token regardless of whether the reasoning trace helps, you are paying nearly three times as much per unit of work as the cohort baseline.
Two more things complicate the "notably fast" framing that Artificial Analysis itself pairs with "very verbose" [13]. Output speed is 154.4 tokens per second against a 104.0 median [5], but time to first token is 168.75 seconds against a median of 1.92 seconds [6], roughly 88 times longer [7]. Nearly three minutes of silence precedes a fast stream. And 130M tokens at 154.4 tokens per second is about 234 hours, or 9.7 days, of serial generation [8] before any parallelism.
Read the price-tier comparison carefully. Artificial Analysis bands proprietary models against both proprietary and open-weights models in the same price range, using a blended 3:1 input/output ratio [12]; at these prices that blend is $0.45 per 1M tokens, landing Luna in the $0.15 to $1 band [9]. The median of 17 is the middle of that band, not the frontier. Separately, the site's headline blended rate of $0.17 per 1M tokens assumes a 7:2:1 cache-hit/input/output mix [11] - an assumption a model this output-heavy will not honour on most real traffic.
Worth watching: whether OpenAI exposes a reasoning-effort control that cuts the 130M figure without moving the 52, and whether the 168.75 second time to first token is a cold-start artefact or the steady state. Also confirm the 1M token context window [7] against your own long-input costs, since input is the cheap half here [3].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
GPT-5.6 Luna (max) scores 52 on the Artificial Analysis Intelligence Index, well above average among other reasoning models in a similar price tier (median: 17).
When evaluated on the Intelligence Index, GPT-5.6 Luna (max) generated 130M output tokens, at the higher end compared to other reasoning models in a similar price tier (median: 60M).
In total, it cost $172.17 to evaluate GPT-5.6 Luna (max) on the Intelligence Index.
GPT-5.6 Luna (max) has a time to first token of 168.75s, at the higher end compared to other reasoning models in a similar price tier (median: 1.92s).
For a blended rate using a 7:2:1 cache hit/input/output ratio, GPT-5.6 Luna (max) prices at $0.17 per 1M tokens.
Artificial Analysis compares proprietary models across proprietary and open-weights models of the same price range, using a blended 3:1 input/output price ratio, with bands of under $0.15, $0.15-$1, and over $1 per 1M tokens.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed but single-source vendor measurements
Every figure is quantitative, specific and internally consistent: the stated output token count and output price reproduce $156.00 of the disclosed $172.17 run cost, and the 3:1 blend of the stated prices lands inside the price band the methodology describes. But it all comes from one publisher, which is also the benchmark operator, with no independent reproduction and no breakdown of the composite score's subtests.
No usage or deployment evidence
The cluster establishes only that the model exists on OpenAI's API and that one benchmark vendor ran it. There is no usage disclosure, customer deployment, traffic, integration or third-party evaluation in the supplied material, so adoption cannot be scored without inventing facts.
Verdict language runs ahead of the cost and latency shape
The numbers behind the story hold up, but the vendor's own verdict language leans favourable relative to what its metrics imply. 'Well priced' rests on a 7:2:1 blended rate of $0.17 while list output price is above the cohort median and verbosity is 2.17 times median, and 'notably fast' sits next to a 168.75s time to first token, about 88 times the cohort median, which is filed as a metric rather than a caveat. Capability-normalised cost is genuinely near cohort, so the gap is framing, not fabrication.
Benchmark operator publishing on its own index and pricing funnel
The only publisher is the operator of the index being cited; the page brands the score, the token-use and cost charts and the price-banding methodology as its own products, and points readers to its provider pricing comparison. That is a visible interest in the index's authority and traffic. The model vendor, OpenAI, is the subject rather than the publisher, and no sponsorship or commercial relationship is disclosed either way in the supplied material.
Arithmetic solid, corroboration absent
Confidence in the derived economics is high because they follow directly from figures the source publishes and reconcile against its stated total cost. Confidence in the wider claim that this is a leading, well-priced model is much lower: one publisher, one harness, no independent replication, no adoption evidence, and cohort medians whose membership is defined by the same publisher.
build
Three frontier launches in a day, all pitched on price. Open weights set the ceiling.4 distinct publishers
leadership
The AI bill nobody reconciles: cost per finished task, not per million tokens1 distinct publisher
product
A 27B laptop model scores like a rented one, and thinks three times as hard to do it1 distinct publisher
build
Grok 4.6 lands in Copilot two days after launch, and the model picker becomes a procurement problem1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 19, 2026