Science1 distinct publisher3 min readPublished
Artificial Analysis priced the new model at $10 and $50 per million tokens, and the threefold token cut it measured in the Codex agent harness covers that increase where the roughly 10% cut on its intelligence suite does not.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
Start with the break-even, because it is a single number. A rate card that rises 2.5x is neutral only if the same job burns 40% or less of the tokens it used to [1]. Everything in the Artificial Analysis writeup sorts around that threshold.
The Codex coding-agent runs clear it. A third of the predecessor's tokens at 2.5x the price works out to about 0.83 of the old bill, roughly 17% cheaper per task [2], and Artificial Analysis measures cost per task as about level with GPT-5.6 Sol at max effort, with two more index points [6]. The gap between that arithmetic and their measurement is the part a token ratio hides: input, cached input and output are priced separately, and the mix moves.
The Intelligence Index does not clear it, and the way it misses is instructive. Output tokens fall about 10% at max effort [8]. Price 90% of the old output volume at 2.5x and you would expect about 2.25x the cost, or 125% more; Artificial Analysis reports 75% more [9], which implies the blended token bill actually fell to roughly 70% of the predecessor's [3]. The efficiency gain is real and larger than the output-token figure alone suggests, but it does not reach 40%.
One caution about the coding scores. Astra's 67 was measured in Codex, Fable 5 in Claude Code, and Muse Spark 1.3 in Muse Code [3], with Fable 5.1 in Claude Code leading on 70 [4]. An index row is a model and a scaffold together, so the threefold token reduction belongs to that pair rather than to the weights alone, and a team swapping models inside its own harness does not automatically inherit it.
The result worth the most attention is the hallucination one. On AA-Omniscience the max-effort hallucination rate drops from 92% to 51%, a 45% relative reduction [5], and accuracy rises 4 points at the same time [10], which rules out the usual trade where a model lowers its error rate by declining to answer more often. What it does not tell you is whether the remaining half lands in your domain.
Elsewhere the ledger is mixed. AA-Briefcase, which runs multi-week projects across thousands of source files, gains about 80 Elo on analytical quality while presentation quality falls, with GPT-5.6 Sol still leading that sub-score [11]. GDPval-AA v2, adapted from OpenAI's dataset covering 44 occupations, drops about 80 Elo [13]. Those two 80s sit on different scales and do not cancel. Humanity's Last Exam gains 6 points [12], while customer support, scientific Python and long-context reasoning each slip 2 to 3 points [14].
The read I would commit to, conditioned on these harnesses: the upgrade pays for itself where an agent loop is the token consumer and the score holds flat or slightly better. On the intelligence side, an equal score of 61 against Fable 5.1's 66 [7][4] is now bought at a higher rate, which makes that purchase a decision about hallucination rate rather than about capability.
Ranked by verification strength, evidence, and original report placement.
GPT-6 Astra's pricing is 2.5x GPT-5.6 Sol's current prices across the board, up from $4/$20 to $10/$50 per million input/output tokens.
GPT-6 Astra keeps the same 90% discount for cache reads and 25% premium for cache writes as its predecessor.
In Codex, GPT-6 Astra scores 67 on the Artificial Analysis Coding Agent Index, approximately equal to Claude Opus 5 and Fable 5 in Claude Code, and Muse Spark 1.3 in Muse Code.
Fable 5.1 in Claude Code leads the Coding Agent Index with a score of 70.
GPT-6 Astra uses one third of the tokens of GPT-5.6 Sol (max) in the Codex harness and one fifth of the tokens of Claude Opus 5 (xhigh); Artificial Analysis describes this as 70% more token efficient than GPT-5.6 Sol.
At max effort, GPT-6 Astra costs about the same per task as GPT-5.6 Sol (max) while scoring 2 points higher on the Coding Agent Index, and per task is less than half the cost of Claude Fable 5 for the same score.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 4, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Spark 1.3's index jump lands on the three tests that carry half the score6 distinct publishers
science
Frontier scores at $2/$6: Grok 4.6 ties GPT-5.6 on one evaluator's index for a fifth the output price1 distinct publisher
build
Developer habit, priced at $965B: what Anthropic's run actually proves1 distinct publisher
leadership
Meta's Muse Spark 1.3 matches three flagship models at 55 cents a task1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Granular, and all from one ruler
Artificial Analysis is the primary source for what it measured, and it measures at a level of detail most launch coverage never reaches: effort settings named, harnesses named, index version pinned at 4.1.1, regressions itemised rather than glossed. The ceiling is that nobody has held a second ruler to any of it — no replication, no provider confirmation of the prices being multiplied, and the per-task cost figures rest on an input, output and cache mix the post does not break out.
Priced, not yet observed in use
A rate card and a benchmark sweep tell you the model is buyable and how it scores; neither tells you who is running it. There is no deployment, no customer, no traffic figure and no disclosed usage anywhere in this reporting, so we leave uptake unscored rather than let a launch-week evaluation stand in for traction.
Candid on scope, loose on the arithmetic
Credit where it is due: the same post that headlines a coding-harness cost win also prints the 75% per-task premium on the intelligence suite, the roughly 80-Elo fall on GDPval-AA v2 and three smaller regressions. The stretch is quieter and numerical. A cut to one third of the tokens at 2.5x the rate should come out near 17% cheaper, not 'about the same', and a 10% output-token saving at that rate should imply 125% more per task, not 75% — so unshown mix effects are doing work in both directions. Add that 'hallucinates half as much' still means wrong on roughly half the items, and the frame runs a little ahead of the numbers.
The scorekeeper is the only witness
Every quantity a reader takes away is denominated in something Artificial Analysis owns and brands — Coding Agent Index, Intelligence Index v4.1.1, AA-Omniscience, AA-Briefcase, its own adaptation of GDPval — and the piece closes by inviting you to compare models on its site. That is an interest in index salience rather than in any particular model winning, and the willingness to publish regressions cuts against favouritism. Still, the model's vendor is absent from the story, the price rise is described only by the party whose product is the comparison, and no commercial relationship is addressed in either direction.
Trust the measurements, hold the framing loosely
The individual numbers are about as reliable as single-lab numbers get, and they are internally consistent enough that the derived arithmetic only strains at the edges. What we cannot yet stand behind is the headline proposition that efficiency absorbs the price rise: it is true inside one harness at one effort setting, on one firm's cost model, with no usage data and no second measurement to test whether real workloads look like the Codex runs.