Build2 distinct publishers3 min readPublished
Artificial Analysis scores the same model level with its predecessor, and OpenAI charges two and a half times as much per token, so the ranking you inherit is a claim about a test mix that is not yours.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Epoch's own per-benchmark table explains the split. Astra leads on math, knowledge, and puzzles, Fable 5.1 holds the top mark on nearly every coding test, and Epoch has logged exactly one coding score for Astra, from a run at medium reasoning effort [10]. Artificial Analysis weights knowledge, coding, and text comprehension [3]. A composite is a weighting of somebody else's test mix, and these two do not share one.
The per-token sticker is the wrong unit to decide on. According to the-decoder, OpenAI charges two and a half times as much per unit of processed text as it did for Sol, which works out to roughly 75 percent more per task [4]. Those two figures pin the token count: 1.75 divided by 2.5 is 0.70, so a task now consumes about 70 percent of the text it did on Sol [1]. The saving comes from step count. Astra needs a third of Sol's compute steps and a fifth of Opus 5's [5], and on the Coding Agent Index it scores 67 at roughly a third of Sol's token usage, behind Fable 5.1 at 70 [7].
So the 2.5x transfers in full only to work with no steps to save. A single-shot long-context job absorbs the whole increase, and long-context reasoning is one of the places Astra slipped, along with banking support, SciCode, and about 80 Elo on GDPval-AA v2 [9]. An agent that loops many times per ticket keeps most of the discount: on coding, Artificial Analysis has Astra level with Claude Fable 5 at less than half the cost per task [6].
ARC-AGI-3 shows the same mechanism at its extreme. On the standard scaffold the run cost $49,791 with reasoning off and scored 35.2 percent; at maximum reasoning it cost $26,098 and scored 62.7 percent [16]. That is $1,414 per point against $416, a 3.4x improvement in cost per point from turning the dial up [2]. ARC Prize attributes it to Astra finishing the games in fewer moves, which means fewer model calls and fewer tokens [16]. One setting breaks the pattern. The "low" level scores 17.5 percent, below running with no reasoning at all, and ARC Prize offers no explanation [17]. A dial whose second notch is worse than off is one you measure per workload rather than tune by intuition.
The 99.9 percent OpenAI reported came from its own harness, which keeps reasoning chains alive between requests and automatically summarizes long runs [13]. ARC Prize measured those runs at about 3.66 times faster with 49 percent fewer tokens, across 167 game-reasoning pairs both setups solved [14]. Both features are client-side design choices, so the gap is partly reproducible, but 49 percent is the saving ARC observed on one benchmark's shared subset rather than a number you inherit by bolting on a summarizer.
Epoch applies the same discipline to its Erdos result. Astra was the only model to close two of 68 open problems with Lean-verified proofs on a budget of $300 per attempt, and three further solutions from non-standardized runs costing more than $220,000 were excluded from the score [20]. That excluded budget is more than 700 times the standardized per-attempt allowance [3]. Chollet's forecast revision rests on the efficiency side of the same benchmark, where Astra beat average human efficiency for the first time and he called the progress twice as fast as he expected [18][19]. The measurement that transfers is the one you run on your own mix, at the effort level you would actually ship.
Ranked by verification strength, evidence, and original report placement.
Epoch AI combines more than 50 benchmarks into one overall score and puts GPT-6 Astra clearly in first place with 169 points, ahead of 267 models.
Artificial Analysis rates GPT-6 Astra at 61 points, exactly level with its predecessor and behind Claude Fable 5.1 at 66 points.
Artificial Analysis tests knowledge, coding, and text comprehension.
ARC-AGI-3 drops an AI into unfamiliar game worlds whose rules and goals nobody explains, and the model has to work out what to do by trial and error.
GPT-6 Astra reaches 62.7 percent on ARC-AGI-3 at a test cost of roughly $26,000; predecessor GPT-5.6 Sol managed 7.78 percent and Claude Opus 5 got 30.16 percent, while Fable 5 and Fable 5.1 are not on the benchmark yet.
On ARC-AGI-3, GPT-6 Astra works more efficiently than the average human for the first time.
Distinct publishers with included, body-backed reporting in this cluster.
2 articles · September 4, 2026
4 articles · September 4, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
science
Token efficiency absorbs GPT-6 Astra's 2.5x price increase inside the coding harness1 distinct publisher
invest
Astra's 99.9% holds up only on the harness OpenAI ran itself1 distinct publisher
invest
Anthropic Cuts Cache-Read Prices by 75%; Cache Reads Were ~60% of a Heavy Agent's Bill Before the Cut2 distinct publishers
build
Three frontier launches in a day, all pitched on price. Open weights set the ceiling.4 distinct publishers
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Named scorers, nobody re-running
The figures here are unusually well-attributed for a launch story: Epoch AI, Artificial Analysis and ARC Prize are each named with their methodology attached, and ARC Prize even publishes the harness comparison that undercuts the vendor number. What is missing is any second measurement. Both outlets sit downstream of the same four sources, and the clearest sign of it is that they print 99.9 and 98.6 percent for what is described as the same vendor-harness run without either noticing the other figure exists.
Scoreboards and a price list
Astra is shipped and metered — there is a per-token rate, a predecessor it supersedes, and vendor demonstrations inside desktop software. Nothing in this reporting describes anyone actually running it: no customer, no workload, no volume, no deployment. Grading uptake from a leaderboard position would be inventing the part of the story that has not happened yet.
First place with one coding run behind it
The travelling claims are 'ranked first' and 'aced the hardest benchmark'. Underneath the first, Epoch has logged a single coding score for Astra at medium reasoning while Fable 5.1 tops nearly every coding test, and a second scorer has the model exactly level with its predecessor. Underneath the second, the vendor-comparable figure is 62.7 percent, not the near-perfect run OpenAI's harness produced. Epoch's own handling of the Erdős results shows the discipline the headlines lack: two solutions counted at $300 an attempt, three discarded because they cost more than $220,000.
No disinterested party in the room
OpenAI published the near-perfect figure produced on its own harness after a version of that dispute had already occurred over Sol, and is charging two and a half times more per token for the model it describes. ARC Prize's chief pulls his AGI forecast forward on the strength of his organisation's own benchmark. Epoch AI and Artificial Analysis both sell composite rankings and arrive at opposite verdicts from the same model. None of that implies bad faith, and Epoch's exclusion of the $220,000 runs cuts against its own headline; it does mean every number in this story reaches the reader through someone with a position in it.
Firm on the figures, loose on their meaning
Individual numbers can be traced and checked, and the disagreements are documented rather than hidden, so we hold the specifics with reasonable confidence. What we cannot yet stand behind is any statement about what Astra is worth in production: two scorers contradict each other, the coding evidence is one run deep, the model's efficiency claim depends on which harness ran it, and the 'low' reasoning setting scoring below no reasoning at all is an anomaly nobody has explained.