Security1 distinct publisher2 min readPublished
Astra scored a perfect 100% on OpenAI's own exploit-development benchmark, against 78.5% for GPT-5.6 Sol, and the shipped model's refusal to write proof-of-concept code is a policy the company has already said it will relax.
The Watch · Security desk

Compiled by The WatchSomething wrong?How this is made
Read the two ExploitBench numbers as failure rates. GPT-5.6 Sol missed 21.5% of the tasks, about one in five. Astra missed none [16]. The benchmark scores the step between an advisory and working code, turning known software flaws into exploits [3], so what matters to a patch queue is less the score than what the score was measured against.
The announcement does not say. Task count, target software, an evaluation run by anyone other than OpenAI, a jailbreak success rate, any observed misuse: none of it appears, and every capability figure is the company's own reporting [19].
The zero-day claims sit apart from the benchmark. OpenAI says Astra reaches substantially higher arbitrary code-execution rates than Sol against flaws it dates as the previous three months, July to August 2026, and that the set included two zero-days in software it does not name [6]. July to August is two months [18]. Small, but it is the only description offered of how fresh the tested bugs were.
The brake is server-side, and it has a date on it. OpenAI says less restrictive safeguards arrive through Daybreak in the coming weeks, opening up vulnerability and proof-of-concept validation, malware analysis, and detection engineering [9]. The company frames the whole release around what it calls a defender's window, a narrowing opportunity to close gaps before attackers seize them [15]. It also says it strengthened the model against jailbreaks, gave its monitoring systems more context, and added safeguards to detect and contain misalignment [11].
For the rest of the stack, the numbers that travel are FrontierMath Tier 4 at 98% and ARC-AGI-3 at 99.9% [5]. Those are the ones a procurement deck will quote. ExploitBench is the one a vulnerability manager should read, because a perfect score on advisory-to-exploit conversion is a statement about tempo, and tempo is what remediation SLAs are actually built on.
On OpenAI's own account the weights crossed the Critical cybersecurity threshold in its Preparedness Framework before the model shipped [2]. The restraint that came with them is configuration, applied per request, and scheduled by the vendor to relax. Anyone whose patch windows assume public exploit code trails an advisory by days should test that assumption against the tier OpenAI has promised for the coming weeks rather than the one available on launch day.
Ranked by verification strength, evidence, and original report placement.
OpenAI said Astra saturates ExploitBench with a 100% score; ExploitBench evaluates a model's ability to turn known software vulnerabilities into working exploits.
GPT-5.6 Sol, OpenAI's previous frontier cyber-capable model, scored 78.5% on ExploitBench, against Astra's 100%.
OpenAI on Thursday officially unveiled GPT-6 Astra, describing it as the "world's most intelligent and aligned model."
Days before the launch, OpenAI said GPT-6 Astra had reached the "Critical" cybersecurity capability threshold under its Preparedness Framework.
OpenAI reported that Astra saturates FrontierMath Tier 4 with a 98% score and ARC-AGI-3 with a 99.9% score.
OpenAI said Astra achieves substantially higher arbitrary code-execution rates than GPT-5.6 Sol when tested on flaws from "the previous three months between July and August 2026," a set that included two zero-day vulnerabilities in unspecified software.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 3, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
product
OpenAI's Astra turns the unattended desktop session into the thing teams delegate4 distinct publishers
invest
OpenAI routes its first Critical cyber model to market through an alpha allowlist1 distinct publisher
product
OpenAI ships a computer-use agent it classifies as a critical cybersecurity capability6 distinct publishers
product
OpenAI gates its first 'critical' cyber model behind an early-access partner list1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One vendor's numbers, cleanly relayed
Nothing in this story has been checked by anyone outside OpenAI. The 100% and the 78.5% arrive without a task set, a task count or a grading method; the jailbreak hardening is asserted without a success rate; the two zero-days are in software nobody names. The Hacker News reports the announcement accurately, and accuracy about a claim is not evidence for it. The date slip — three months that run July to August — is small, but it is the one place where the announcement can be checked against itself, and it fails.
A first cohort and one named pilot
Real but thin. Astra is with a small, unnamed set of organizations; the broad paid-tier and cloud availability is stated as expectation, not fact. On the defender side, one partner is named — MS-ISAC, for public-sector and water-system teams — and the $1 billion behind Daybreak is a commitment, not spend. No seat counts, no customer names, no usage disclosed.
Saturation language outruns what is shown
"World's most intelligent and aligned model," three benchmarks saturated, a perfect exploit-development score — and the entire supporting record is the announcement making those claims. The gap widens because the safety story is offered as reassurance while being marked temporary in the same breath: refusals today, "less restrictive safeguards" in weeks. The one deployment fact strong enough to justify the alarm, distribution through Azure and Bedrock, gets less space than the superlatives.
The seller sets the test and funds the buyers
OpenAI declares its own Critical threshold, reports its own score on the exploit benchmark, and frames the moment as a "defender's window" that is closing — urgency that happens to point at the product. Then it commits $1 billion to subsidize access for water utilities, grid operators, banks and local government, which is philanthropy and pipeline at once. The Hacker News has no stake in the claims but does have an audience for which a Critical-rated cyber model is irresistible, and the story is built almost entirely from the company's release.
Provenance certain, substance unaudited
We are confident about what was said, by whom, and when: the sourcing chain is short and completely visible. We are not confident about whether any of it holds. That asymmetry is stable rather than fragile — an independent ExploitBench run or a third-party red-team result would move the capability read sharply without changing anything about the announcement record itself.