Invest1 distinct publisher3 min readPublished
Astra scored 98.6% on ARC-AGI-3 where its predecessor managed 7.8%. OpenAI is still rationing access while it scales capacity, and that tells a clerical-automation budget more than the benchmark does.
The Investor · Invest desk

Compiled by The InvestorSomething wrong?How this is made
A vendor that rations is saying something about cost, and OpenAI's stated reason for holding Astra back at first is that it needs to scale capacity for what a company spokesperson called "a very large model" [11]. That is the most useful sentence in the launch for anyone drafting an automation budget, because it prices the capability in the currency that is actually scarce, serving capacity rather than model quality, and because there is no published cost per session and no error rate on repetitive clerical work anywhere in the announcement [23].
Read the benchmark spread precisely. Astra's ARC-AGI-3 result sits 90.8 points above the model OpenAI shipped before it, roughly 12.6 times the score [19], and about 3.3 times Anthropic's Claude Opus 5 [20]. ARC-AGI-3 is meant to test reasoning applied to situations a model has not seen before [24], which is nearly the opposite of the work being budgeted: the same purchase-order layout again and again, where the expensive failure is a confidently wrong field that nobody catches until reconciliation. The launch demonstration itself was a person instructing a computer by voice in a video that Fortune described as seamless albeit staged [4].
Notice what a finance team that underwrites next year's clerical headcount on this model is quietly deferring: permissions and audit logs, the unglamorous plumbing that decides whether an agent is allowed near a general ledger at all. Anthropic pioneered computer use with a beta in 2024 and Perplexity shipped a system that navigates a virtual computer in February, and the concept is still not a mainstream part of working on a computer for most people [15]. Two years of availability without adoption is evidence that the binding constraint sat downstream of the model.
If the reason computer use never landed was simply that the models were not good enough, then the 90.8 points are the whole story and the integration objection is a lagging indicator that will look silly by summer. More likely, both hold, and sequencing decides the budget: the capability arrives now, the permissions arrive whenever IT and finance finish arguing about who signs off when an agent submits the wrong form. Brockman told reporters it is not unreasonable to feel we are in the AGI era and reasonable to call this the first such model [16]; OpenAI once defined AGI as a system performing all economically valuable work as well or better than humans [17], which is a labor-cost claim, and labor-cost claims get settled by invoices rather than scores.
A published price per computer-use session alongside an error rate on a named real workflow would make the substitution arithmetic possible for the first time, and it would move me. A general rollout that is not capacity-gated would move me too, telling you OpenAI thinks the marginal session is cheap enough to sell in volume. Until one of those lands, the honest line item is a pilot.
Ranked by verification strength, evidence, and original report placement.
On ARC-AGI-3, Astra scored 98.6%, GPT-5.6 Sol had scored 7.8%, and Anthropic's Claude Opus 5 had scored 30%.
OpenAI submitted Astra to the U.S. government for review ahead of release under the terms of a loosely defined, voluntary AI safety framework agreed between tech companies and the Trump administration whose details have not been made public.
ARC-AGI-3 is a difficult benchmark meant to test a model's ability to apply reasoning to situations it has not encountered before.
OpenAI released GPT-6 Astra, its most powerful model, and one of its hallmark features is called "computer use," in which the model aims to navigate a computer as a human would.
OpenAI co-founder and president Greg Brockman said the model "can zip through spreadsheets, fill out forms, and navigate across webpages often at superhuman speed," and called computer use "a particularly important part of what's new."
OpenAI shared a video with reporters of someone sitting in a chair instructing the computer through their voice in what appeared to be a seamless, albeit staged, interaction.
Distinct publishers with included, body-backed reporting in this cluster.
fortune.com
1 article · September 3, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
product
OpenAI ships a computer-use agent it classifies as a critical cybersecurity capability5 distinct publishers
security
Frontier labs put their best vulnerability-hunting models behind vetted-defender lists3 distinct publishers
invest
OpenAI rates GPT-6 Astra capable of hacking hardened systems without human guidance1 distinct publisher
product
OpenAI prices its own guardrails: 20% more compute, plus a two-week training pause1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One briefing, no outside check
Every number that matters here — 98.6% on ARC-AGI-3, 100% on ExploitGym, more than 100,000 GPUs at Stargate — was released by OpenAI and relayed by Fortune the same day. Nobody outside the company has run the tests, the government review that preceded release is explicitly undisclosed, and the reassurance that Astra had no part in July's escape is the company's own. Fortune's care raises this above the floor: it attributes each claim to a named executive and marks the demo video as staged. The same passage still calls Claude Fable 5.1 Anthropic's best model while scoring Astra against Claude Opus 5, and nobody has cleaned that up.
Gated to one program
Today the model is in the hands of vetted cybersecurity customers in the Daybreak program and no one else. Plus, Pro and Enterprise access, the API and AWS are all "in the coming days" — a phrase doing heavy lifting with no date behind it. There is dated, specific deployment to point at, which is why this is measured rather than blank, but no customer names, no seat counts and not one workflow anybody has run in production.
Claims outrun the rollout
"Superhuman speed" at filling in forms, and a president willing to call this the first model of the AGI era, sit awkwardly beside a release throttled because the model is too large to serve. The tell is what the launch never priced: no cost per task, no error rate on the clerical work it claims to do faster than people, and a demonstration OpenAI staged. Fortune supplies its own corrective — don't expect this in every office overnight — but the AGI line is still measured against a definition, OpenAI's own, about performing all economically valuable work.
The vendor set the table
This is a launch briefing, and almost every load in it is carried by the party that benefits: the capability claims, the scores, the assertion that Astra was absent from July's breakout, and the judgement that new monitoring suffices. Fortune's counterweights are two and both thin — unnamed skeptics on the monitoring, and Brockman declining to describe the government process when pressed on it. The staged video and the undisclosed federal review point the same way: the company chose what could be seen.
Firm on what was said
We are on solid ground about who said what: Fortune names Brockman, Clark and Glease, quotes them directly, and marks the staged demo rather than passing it off. We are on much softer ground about whether any of it holds — one outlet, one controlled briefing, no replication, and the two facts a buyer would weigh first, price and reliability, never entered the record.