Build1 distinct publisher3 min readPublished
Artificial Analysis scores the same model level with its predecessor, and OpenAI charges two and a half times as much per token, so the ranking you inherit is a claim about a test mix that is not yours.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Epoch's own per-benchmark table explains the split. Astra leads on math, knowledge, and puzzles, Fable 5.1 holds the top mark on nearly every coding test, and Epoch has logged exactly one coding score for Astra, from a run at medium reasoning effort [10]. Artificial Analysis weights knowledge, coding, and text comprehension [3]. A composite is a weighting of somebody else's test mix, and these two do not share one.
The per-token sticker is the wrong unit to decide on. According to the-decoder, OpenAI charges two and a half times as much per unit of processed text as it did for Sol, which works out to roughly 75 percent more per task [4]. Those two figures pin the token count: 1.75 divided by 2.5 is 0.70, so a task now consumes about 70 percent of the text it did on Sol [1]. The saving comes from step count. Astra needs a third of Sol's compute steps and a fifth of Opus 5's [5], and on the Coding Agent Index it scores 67 at roughly a third of Sol's token usage, behind Fable 5.1 at 70 [7].
So the 2.5x transfers in full only to work with no steps to save. A single-shot long-context job absorbs the whole increase, and long-context reasoning is one of the places Astra slipped, along with banking support, SciCode, and about 80 Elo on GDPval-AA v2 [9]. An agent that loops many times per ticket keeps most of the discount: on coding, Artificial Analysis has Astra level with Claude Fable 5 at less than half the cost per task [6].
ARC-AGI-3 shows the same mechanism at its extreme. On the standard scaffold the run cost $49,791 with reasoning off and scored 35.2 percent; at maximum reasoning it cost $26,098 and scored 62.7 percent [16]. That is $1,414 per point against $416, a 3.4x improvement in cost per point from turning the dial up [2]. ARC Prize attributes it to Astra finishing the games in fewer moves, which means fewer model calls and fewer tokens [16]. One setting breaks the pattern. The "low" level scores 17.5 percent, below running with no reasoning at all, and ARC Prize offers no explanation [17]. A dial whose second notch is worse than off is one you measure per workload rather than tune by intuition.
The 99.9 percent OpenAI reported came from its own harness, which keeps reasoning chains alive between requests and automatically summarizes long runs [13]. ARC Prize measured those runs at about 3.66 times faster with 49 percent fewer tokens, across 167 game-reasoning pairs both setups solved [14]. Both features are client-side design choices, so the gap is partly reproducible, but 49 percent is the saving ARC observed on one benchmark's shared subset rather than a number you inherit by bolting on a summarizer.
Epoch applies the same discipline to its Erdos result. Astra was the only model to close two of 68 open problems with Lean-verified proofs on a budget of $300 per attempt, and three further solutions from non-standardized runs costing more than $220,000 were excluded from the score [20]. That excluded budget is more than 700 times the standardized per-attempt allowance [3]. Chollet's forecast revision rests on the efficiency side of the same benchmark, where Astra beat average human efficiency for the first time and he called the progress twice as fast as he expected [18][19]. The measurement that transfers is the one you run on your own mix, at the effort level you would actually ship.
Ranked by verification strength, evidence, and original report placement.
Epoch AI combines more than 50 benchmarks into one overall score and puts GPT-6 Astra clearly in first place with 169 points, ahead of 267 models.
Artificial Analysis rates GPT-6 Astra at 61 points, exactly level with its predecessor and behind Claude Fable 5.1 at 66 points.
Artificial Analysis tests knowledge, coding, and text comprehension.
OpenAI charges two and a half times as much per unit of processed text for GPT-6 Astra, which makes a task cost roughly 75 percent more than it did with GPT-5.6 Sol.
GPT-6 Astra needs only a third of the compute steps Sol uses and a fifth of what Opus 5 uses.
On coding tasks Astra hits the same score as Claude Fable 5 according to Artificial Analysis, but costs less than half as much per task.
Distinct publishers with included, body-backed reporting in this cluster.
Follow any of these and your For You feed starts watching them — no settings page required.
leadership
ARC Prize puts Astra 37 points below the score OpenAI led with3 distinct publishers
invest
Anthropic Cuts Cache-Read Prices by 75%; Cache Reads Were ~60% of a Heavy Agent's Bill Before the Cut2 distinct publishers
invest
Compute scarcity meters the model OpenAI says can fill out forms at superhuman speed1 distinct publisher
build
Three frontier launches in a day, all pitched on price. Open weights set the ceiling.4 distinct publishers
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Named evaluators, single relay
The numbers themselves are unusually specific and each is pinned to the lab that produced it — Epoch AI, Artificial Analysis, ARC Prize. What is missing is a second pair of eyes: nothing in our coverage links to or quotes the underlying evaluation write-ups, and the most consequential figures, the 62.7 versus 99.9 percent split and the single medium-reasoning coding entry, are read out of dashboards we have not seen.
Scoreboards, not deployments
Four separate evaluation efforts have already run this model, which tells you the benchmarking market moves fast — it tells you nothing about who is shipping on it. The only commercial fact in the reporting is the price, and the only usage figures are test harness token counts.
The headline number is the vendor's
The distance between 99.9 percent and 62.7 percent is the whole gap in one line, and it exists because OpenAI's scaffold, not the model alone, did part of the work. The Decoder deflates it rather than amplifying it — that is why the score is modest rather than large — but the claims now in circulation are a first-place ranking resting on one coding datapoint, a human-efficiency crossover measured on the vendor harness, and an AGI forecast moved forward on the back of it.
Everyone here scores something
OpenAI supplied the harness that turned a 62.7 into a 99.9, and this is the second generation running that the same disagreement has surfaced. Epoch AI and Artificial Analysis compete to be the scoreboard people cite, and they reached opposite verdicts on the same model. ARC Prize both runs the test and employs the man revising his AGI timeline upward on its results. None of that makes the numbers wrong; it does mean no one in the story is a disinterested party.
Coherent, but uncorroborated
Internally the account holds together: prices, token ratios and cost-per-point all reconcile when you do the arithmetic, and the caveats are placed next to the claims they undercut rather than at the bottom. What keeps this in the middle is that one publisher is the only route to all of it, and the two loose threads — the 'low' reasoning setting scoring below no reasoning, and Epoch's near-empty coding column — are left unexplained by the evaluators themselves.