Leadership3 distinct publishers3 min readPublished
The same model produced 99.9% in OpenAI's launch post and 62.7% on the benchmark authors' neutral harness, and Astra's input tokens cost double GPT-5.6 Sol's, which leaves the vendor table doing very little work in a purchase decision.
The Board Room · Leadership desk
Compiled by The Board RoomSomething wrong?How this is made
The two figures describe the same model on the same benchmark run through different plumbing. ARC Prize's Standard harness hands every model an identical minimal, provider-neutral interface and makes it decide for itself what to keep in visible notes [9]. The Provider Adapter preserves OpenAI's opaque reasoning state between requests and compacts longer conversations, and across the game-and-reasoning pairs both setups solved, adapter runs used 49% fewer tokens and finished about 3.66 times faster [10][11]. The distance between the two results is a description of the scaffold as much as of the model.
The cost line inverts the usual assumption that the better score is the expensive one. Astra at maximum effort scored 62.7% on the semi-private set for $26,098, while the high-effort Provider Adapter run that produced 99.9% cost $18,817 [12], so the lower score cost about 39% more [13]. Persistent reasoning state made the model cheaper and stronger at the same time, an engineering result rather than a sleight of hand. The marketing is in the comparison, not the adapter.
The adapter is how paying customers will actually call the model, so the neutral figure understates what a buyer receives. ARC Prize agrees on direction, calling Astra "a noticeable step-function change in frontier model capabilities" and "a major milestone worth celebrating" [14]. The same evaluation says saturating ARC-AGI-3 would not prove AGI, and that the benchmark's mechanics are deterministic and closed-ended and do not represent the complexity of the real world [15]. That gap between a capability signal and a procurement input is exactly what a buyer has to bridge on their own.
Pricing is where the buyer's exposure gets concrete. Sol Standard runs $5 per million input tokens and $30 per million output [7], so a task consuming a million tokens each way costs $35 on Sol against $60 on Astra, a rise of about 71% [28], and Astra must finish the same work with roughly 42% fewer tokens to reach parity [17]. OpenAI says the model can do exactly that by spending fewer tokens per task [16], a claim only a buyer's own traffic can settle. It is worth crediting that OpenAI published a losing number too: its own table puts Astra at 57.2% on Humanity's Last Exam with tools against Sol's 65.0% [6].
The absent benchmark is the one that matches the claim being made. OpenAI's charter defines AGI as "highly autonomous systems that outperform humans at most economically valuable work" [18], and GDPval is its own benchmark for economically valuable work [8]. Days before the launch, Altman told the Sources podcast that AGI is "at best a very poorly defined term" and "like an irrelevant marketing term" [19]. The Guardian notes the company is pushing toward a stock market listing it hopes will value it above $850bn [26], which is the visible incentive behind a launch-day label.
What a buyer decides this quarter is narrow, because the rollout is narrow. Astra went first to a limited set of organizations including OpenAI's Daybreak program, with ChatGPT tiers, the API and Amazon Bedrock following over days, and enterprise access stays off until an administrator turns it on [20][21]. Because the model carries OpenAI's critical cybersecurity classification, less restrictive access sits with an initial set of trusted defenders [22], and chief scientist Jakub Pachocki said confidence in monitoring may constrain further development [23]. So the trade-off is legible: a buyer either pays about 71% more per equal-token task now for a gain the vendor's headline number cannot evidence, or waits a quarter, measures token consumption on its own workloads, and buys the same model on numbers it owns.
Ranked by verification strength, evidence, and original report placement.
OpenAI released GPT-6 Astra on September 3, 2026, and president Greg Brockman closed the press briefing by saying "Welcome to the AGI era."
OpenAI led its launch case with a 99.9% score on ARC-AGI-3, a test of how agents learn unfamiliar interactive environments.
ARC Prize, which built the benchmark, scored the same model at 62.7% on its provider-neutral Standard harness and said it is not claiming Astra is AGI.
OpenAI's standing company charter defines AGI as "highly autonomous systems that outperform humans at most economically valuable work".
Days before the launch, OpenAI chief executive Sam Altman told the Sources podcast that AGI is "at best a very poorly defined term. I was going to say it's like an irrelevant marketing term."
Astra has what OpenAI calls a "critical" level of cybersecurity capability, and the company will allow less restrictive access only to an initial set of trusted cybersecurity defenders.
Distinct publishers with included, body-backed reporting in this cluster.
businessinsider.com
1 article · September 3, 2026
implicator.ai
1 article · September 3, 2026
theguardian.com
2 articles · September 3, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
product
OpenAI ships a computer-use agent it classifies as a critical cybersecurity capability5 distinct publishers
invest
OpenAI rates GPT-6 Astra capable of hacking hardened systems without human guidance1 distinct publisher
product
OpenAI prices its own guardrails: 20% more compute, plus a two-week training pause1 distinct publisher
invest
Compute scarcity meters the model OpenAI says can fill out forms at superhuman speed1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Precise numbers, one outside scorer
The specific figures in this story are unusually well pinned down — two scores, two run costs, two token prices, a named harness for each — and ARC Prize's own caveats are quoted rather than paraphrased. What thins the record is provenance: apart from ARC Prize and Artificial Analysis, nearly every capability figure is OpenAI measuring its own model in its own environment, and even ARC Prize's evaluation reaches readers here through Implicator.ai's account of it.
Announced everywhere, deployed narrowly
One day in, Astra exists in production for a limited set of organizations and Daybreak enterprise customers; paid ChatGPT tiers, the API and Amazon Bedrock are promised rather than live, and enterprise admins have to flip it on themselves. Nobody in this reporting cites a user count, a seat count, a workload, or a customer naming a task it has moved to Astra. The only usage figures on the record are benchmark runs.
The label outruns the scoreboard
A briefing that closes on "Welcome to the AGI era" is measured against a company definition about outperforming humans at most economically valuable work — and the benchmark OpenAI built for exactly that, GDPval, is missing from the launch record. The headline 99.9% halves to 62.7% on the benchmark authors' own neutral interface, the authors explicitly decline the AGI label, an outside index puts Astra level with the model it replaces, and OpenAI's own table has it below Sol on Humanity's Last Exam. The gap is not fabrication; it is a real step forward wearing a much larger word, and the company's own chief executive had called that word an irrelevant marketing term days earlier.
A launch inside an IPO window
The Guardian supplies the motive that the launch coverage otherwise floats free of: OpenAI is chasing a listing valued above $850bn while Anthropic eyes one of its own, and more than a thousand frontier-lab employees have warned in writing about competitive pressure not to slow down. Every capability figure except two comes from the seller, and the selected figure is the one from the harness that preserves the seller's own hidden state. Worth noting the counter-pressure too: the outside scorer is also the benchmark's author, so its 'step-function change' language sits close to its own interest in ARC-AGI-3 remaining the yardstick.
Load rests on a single outlet
The launch events are triple-reported and consistent, so the framing is solid. The analytical spine is not: the harness gap, the run costs, the token rates and the GDPval omission all come from Implicator.ai alone, and The Guardian's report is present twice, so our coverage looks broader than the independent reporting behind it. None of the primary documents — ARC Prize's evaluation, OpenAI's launch table, Artificial Analysis's index — is in hand directly.