Science1 publisher3 min readPublished
GPT-5.6 Luna scores 52 against a peer median of 17. Its token count is what lands on your bill
Artificial Analysis puts OpenAI's new reasoning model at three times the index score of its price-tier peers and more than twice their token appetite. The output line dominates the invoice.
The Scientist · Science desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- GPT-5.6 Luna (max) scores 52 on the Artificial Analysis Intelligence Index, well above average among other reasoning models in a similar price tier (median: 17).
- When evaluated on the Intelligence Index, GPT-5.6 Luna (max) generated 130M output tokens, at the higher end compared to other reasoning models in a similar price tier (median: 60M).
- GPT-5.6 Luna (max) costs $0.20 per 1M input tokens (median: $0.25) and $1.20 per 1M output tokens (median: $0.90), based on OpenAI's API.
- In total, it cost $172.17 to evaluate GPT-5.6 Luna (max) on the Intelligence Index.
- GPT-5.6 Luna (max) generates output at 154.4 tokens per second, above the median of 104.0 t/s for reasoning models in a similar price tier.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
Artificial Analysis has published its run on GPT-5.6 Luna (max), the OpenAI reasoning model released on July 9, 2026 [8], scoring it 52 on the Artificial Analysis Intelligence Index against a median of 17 for reasoning models in a similar price tier [1]. The number that will actually show up on an invoice is a different one: the model generated 130M output tokens to complete the index, against a 60M median [2].
That is 2.17 times the token use of the median model in its cohort [1], for 3.06 times the score [11]. Published prices are $0.20 per 1M input tokens and $1.20 per 1M output tokens, against medians of $0.25 and $0.90 [3]. So the input side is cheap and the output side is not, and the output side is where the volume sits. At list price, 130M output tokens cost $156.00 [2]. A model with median verbosity at the median output price would have spent $54.00 on the same evaluation [3], making Luna's output bill 2.89 times larger [4]. Artificial Analysis reports the full evaluation cost at $172.17 [4], which means output tokens are roughly 90.6 percent of it [5]. The remaining $16.17 covers input and cache-hit tokens, which the benchmark does not break out [10].
The honest counterpoint is that verbosity buys something here. Per point of index score, Luna's output spend works out at $3.00 against $3.18 for the median-shaped run, about 5.5 percent cheaper [6]. If your workload is graded on getting the answer right, the extra tokens are not waste. If your workload is fixed-shape, high-volume and priced per token regardless of whether the reasoning trace helps, you are paying nearly three times as much per unit of work as the cohort baseline.
Two more things complicate the "notably fast" framing that Artificial Analysis itself pairs with "very verbose" [13]. Output speed is 154.4 tokens per second against a 104.0 median [5], but time to first token is 168.75 seconds against a median of 1.92 seconds [6], roughly 88 times longer [7]. Nearly three minutes of silence precedes a fast stream. And 130M tokens at 154.4 tokens per second is about 234 hours, or 9.7 days, of serial generation [8] before any parallelism.
Read the price-tier comparison carefully. Artificial Analysis bands proprietary models against both proprietary and open-weights models in the same price range, using a blended 3:1 input/output ratio [12]; at these prices that blend is $0.45 per 1M tokens, landing Luna in the $0.15 to $1 band [9]. The median of 17 is the middle of that band, not the frontier. Separately, the site's headline blended rate of $0.17 per 1M tokens assumes a 7:2:1 cache-hit/input/output mix [11] - an assumption a model this output-heavy will not honour on most real traffic.
Worth watching: whether OpenAI exposes a reasoning-effort control that cuts the 130M figure without moving the 52, and whether the 168.75 second time to first token is a cold-start artefact or the steady state. Also confirm the 1M token context window [7] against your own long-input costs, since input is the cheap half here [3].