Science1 distinct publisher3 min readPublished
OpenAI says Astra can find unknown flaws and chain them into working exploits without a human guiding each step. The evidence published so far is one saturated public benchmark plus a 20-vulnerability internal set, in which the model found two zero-days of its own.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
The result carrying the most weight in OpenAI's update did not come from a leaderboard. In expert-led assessment against a hardened browser, the company says Astra found previously unknown vulnerabilities and built a full compromise chain that escaped the sandbox and executed commands on the host once the browser opened an HTML file [13]. Against a hardened operating system it found multiple flaws and combined them into a local privilege escalation from an unprivileged user to root [14]. Those are existence proofs, and existence proofs are the right instrument for a capability threshold. A benchmark average would not have been.
The benchmark numbers are weaker than they first appear, and OpenAI comes close to saying so. Contamination was its stated reason for building a fresh internal set [6], which is the correct instinct, because a perfect 100% on the public version [5] leaves no headroom at all and the public test can therefore no longer separate Astra from whatever ships next [18]. The replacement holds 20 items, so a single task is worth five percentage points of measured rate [19]. The comparison against GPT-5.6 Sol is reported in words rather than figures: much higher arbitrary code-execution rates, far fewer output tokens, with neither the rates nor the token counts given [7][8].
Token efficiency is the number I would want most, because it is the one that lands on an attacker's invoice rather than an attacker's CV [16]. V8 is also a peculiar proving ground. It is genuinely hardened and heavily fuzzed, which makes success there meaningful, and it carries an unusually deep public literature of exploit write-ups, which leaves generalisation to unfuzzed internal software an open question. The thing this evaluation does not tell you is whether the same chains fall out of codebases with no public exploit corpus behind them.
The safety case is thinner than the capability case, by its nature. OpenAI delayed parts of Astra's development and release for several weeks to harden protections against both misuse and unauthorized model actions [11][20], and it says retrospective testing indicates the safeguards already in production would have prevented the Hugging Face incident, which Astra was not involved in [10]. A retrospective counterfactual, run by the party with an interest in the answer, is the softest evidence in the document even if it is accurate.
There is also daylight between the criteria as written and the evidence as published. The framework's first condition describes functional zero-day exploits of all severity levels across many hardened real-world critical systems without human intervention [3]; the second describes end-to-end novel attack strategies devised and executed from a high-level goal alone [4]. What has been disclosed is one browser, one operating system and one JavaScript engine, with automated benchmarks and experts in the loop [17]. Either the designation is deliberately precautionary against its own wording, or the fuller record is sitting in the system card promised for launch [15]. I would rather a lab over-call a threshold than under-call it, and precaution costs little at this stage. What I cannot do from this document is size the capability, and sizing is what a defender writing next year's budget actually needs.
Ranked by verification strength, evidence, and original report placement.
OpenAI now believes its Astra model meets the Critical cybersecurity capability threshold under its Preparedness Framework, and says it is the first model the company has designated at this level, requiring stronger safeguards during development and before release.
OpenAI describes the Critical designation as meaning that, with the right tools and access, the model can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step.
Under OpenAI's Preparedness Framework, one sufficient condition for the Critical threshold is that the model can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention.
The alternative sufficient condition is that the model can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal; meeting either condition qualifies.
OpenAI ran Astra on ExploitBench, a benchmark evaluating the ability to develop exploits from known vulnerabilities, and the model achieved a perfect score of 100%.
Because of contamination concerns, OpenAI built an internal benchmark called "ExploitBench - Internal Port (June-August 2026)" containing 20 high-severity V8 vulnerabilities disclosed more recently.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 2, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
invest
OpenAI routes its first Critical cyber model to market through an alpha allowlist1 distinct publisher
product
OpenAI gates its first 'critical' cyber model behind an early-access partner list1 distinct publisher
invest
OpenAI allocates Astra's sharpest cyber capability by eligibility instead of price1 distinct publisher
product
OpenAI prices its own guardrails: 20% more compute, plus a two-week training pause1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Two numbers, one publisher
Count the hard figures in this story and you get a 100% and a 20. The rest — code-execution rates, token spend, the identity of the 'hardened' browser and operating system, the two zero-days themselves — arrives as description from the company that ran the tests and wrote the threshold. The exploit chains are specific enough to be checkable in principle, and nobody outside OpenAI is in a position to check them.
Nothing shipped yet
Astra is unreleased. The strongest cyber capability is promised first to an unnamed group of testers, with wider defensive access routed through Daybreak Blue at some later point, and the system card is still to come. The only things that have actually left OpenAI's building are the designation itself and two V8 zero-days now moving toward their maintainers.
The label outruns the arithmetic
'First model ever designated Critical' is an enormous sentence, and what sits under it is a public benchmark Astra just maxed out, a private 20-task replacement where each single vulnerability swings the score five points, and a superiority claim with no numbers attached. The two self-discovered zero-days and the sandbox-to-host chain are genuine, non-trivial work, which is why the gap is moderate rather than wide — it lies between the weight of the designation and what a reader can independently confirm, not between the designation and reality.
Grading its own exam
OpenAI authored the threshold, ran the evaluations, decided the threshold was met, judged its own safeguards adequate to ship, and named its own product as the channel for advanced access. A Critical designation is a safety disclosure and the most emphatic possible capability claim at the same time — it costs the company weeks of delay and invites regulatory attention, but it also makes Astra sound formidable in the same breath that it makes OpenAI sound careful. There is no adversarial party anywhere in that chain.
Sure what was said, not what was shown
What OpenAI stated, and when, is beyond dispute: the post is dated, specific about process, and precise about its own thresholds. Everything downstream of that is a different matter — the capability magnitudes, the safeguard adequacy, the Hugging Face counterfactual and the release timing all rest on internal work nobody else has seen. Two things would move this quickly: maintainer advisories confirming the V8 findings, and the promised system card carrying actual rates.