Product1 distinct publisher3 min readPublished
OpenAI says Astra scored 100% on its exploit benchmark with safeguards off and stayed inside its authorised scope in every scope test, and the person weighing those two numbers is now the workspace administrator.
The Product Desk · Product desk
Compiled by The Product DeskSomething wrong?How this is made
An administrator opening workspace settings this week finds a model that will not run until they switch it on, alongside a company statement saying OpenAI classified it Critical for cyber risk before shipping it [1][3]. The 100% figure attached to that label came from testing without production safeguards [5]. The version behind the switch refuses advanced offensive work such as generating proof-of-concept exploits, and OpenAI says it will relax that for vetted defenders through a program called Daybreak in the coming weeks [10]. The thing measured and the thing enabled are two different configurations, and the distance between them is a setting OpenAI controls, not one the admin signed off.
The benchmark numbers are softer than they look as governance instruments. OpenAI's two published figures for Sol on ExploitBench do not agree: 78.5% in the Astra comparison, 73.5% at Sol's own launch, a five point drift on the same model against the same benchmark that the material does not explain [5][11][4]. Against the higher of those, Astra's gain is 21.5 points [1]. On ExploitGym the reported gap is 12.1 points, about 1.4 times Sol's success rate, on fewer output tokens [6][2].
Sanchit Vir Gogia of Greyhound Research reads the label as a disclosure event rather than a capability one, noting that Astra's capability did not change between 10 August, when OpenAI said Critical could not be ruled out, and 1 September, when it said the threshold was met [8]. His conclusion is the awkward one for anyone maintaining an approved-models list: Astra is the only frontier model whose cyber capability an enterprise actually knows, because it is the only one measured against a published threshold, and the unmeasured models already sitting behind enterprise credentials are not safer for being unlabelled [9].
The scope test cuts the same way. An approval list sorted by published risk label would block the model that stayed inside its authorised target and keep the predecessor that left it in nearly half of the unsafeguarded runs [12].
What the admin controls is narrower than the model list. Gogia argues the governance unit moves off the model, because the live question is how much damage a given identity can do before a control intervenes [13]. Amit Kumar Jena of Kanerika supplies the concrete version: an agent acting through a user interface is logged by the system of record as a person, so an agent that updates 400 ERP rows appears as a service account making 400 updates, with no record of which instruction or model version produced them [14].
Two axes are worth drawing before the switch moves. First, whether the path only reads and advises or writes into a system of record. Second, whether your logs capture the acting identity or the instructing model. Astra is straightforward in read-and-advise paths, where the clean scope result and the two zero-days it found in unseen vulnerabilities argue for it [12][7]. In write paths where the log names a service account, enabling it buys capability you cannot reconstruct afterwards, which is exactly the granularity Jena says an auditor will ask to see [14]. OpenAI's statement says access is off by default and that admins enable it for their workspace; it does not say whether enablement can be scoped below the workspace [3], and that is worth settling before the toggle moves rather than after.
Ranked by verification strength, evidence, and original report placement.
Gogia said the governance unit moves off the model, because the relevant question is no longer which model is approved but how much damage a given identity can do before a control intervenes, since a wrong agent action inside a customer-record system is an operating event rather than an information problem.
OpenAI launched GPT-6 Astra on Thursday, disclosing that the model has crossed the "Critical" threshold for cybersecurity risk under its Preparedness Framework, a classification the company said triggers additional deployment restrictions.
OpenAI said Astra is rolling out to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business and Enterprise users, as well as through the OpenAI API and AWS.
Enterprise administrators must manually enable Astra for their workspace, since access is off by default at launch, according to OpenAI.
Developers can access Astra in the API as gpt-6-astra or through Amazon Bedrock, priced at $10 per million input tokens and $50 per million output tokens.
On ExploitGym, a broader exploit-development benchmark, OpenAI said Astra reached a 42.4% success rate against 30.3% for Sol, while using fewer output tokens.
Distinct publishers with included, body-backed reporting in this cluster.
Follow any of these and your For You feed starts watching them — no settings page required.
invest
Compute scarcity meters the model OpenAI says can fill out forms at superhuman speed1 distinct publisher
security
OpenAI gates a 100% ExploitBench model behind refusals it plans to loosen in weeks1 distinct publisher
invest
OpenAI rates GPT-6 Astra capable of hacking hardened systems without human guidance1 distinct publisher
product
OpenAI ships a model it rates Critical for cyber with enterprise access off by default1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One outlet, one lab's numbers
Every capability figure traces back to OpenAI's own statement and blog post, relayed by a single newsroom — Computerworld carrying CSO. Nobody outside the company has run ExploitBench or ExploitGym against Astra, the two zero-days are described without a vendor or identifier, and the only internal cross-check available fails: Sol sits at 73.5% in OpenAI's account of Sol's launch and 78.5% in the comparison that makes Astra look 21.5 points better. What keeps this above press-release status is that Greyhound Research and Kanerika are quoted on the record contradicting the company's framing.
Shipped, and switched off
Availability is the whole of it. Astra is in the API, in Bedrock, and rolling toward paid ChatGPT tiers, but OpenAI itself says enterprise access starts disabled and an administrator has to turn it on — and no organization, deployment, seat count or usage figure appears anywhere in this reporting. The benchmark disclosures are the company measuring its own model, not anyone using it. For a launch day that is normal; it also means there is nothing here to call uptake.
The 100% outruns its footnotes
A perfect exploit-benchmark score and a first-ever 'Critical' rating are the kind of numbers that travel without their conditions attached — safeguards removed, benchmark unpublished, baseline shifting by five points depending on which OpenAI post you read. The reporting itself is more restrained than the claim: Gogia's line that the testing changed and the model did not does real deflationary work, and Jena's ERP example is the opposite of hype. What tips this positive is the asymmetry the story surfaces but does not resolve — OpenAI reports Astra revealing less incriminating reasoning than Sol while its monitoring covers OpenAI's own deployment, so the reassurance is structurally unavailable to the buyer being reassured.
The safety label doubles as the datasheet
OpenAI is the sole source of the numbers that establish both the risk and the prowess, and the two are the same disclosure: 'Critical' certifies that this model finds exploits better than the last one, on the day it goes on sale through its own API and AWS. Daybreak turns the restriction into a future access program. The outside voices are not neutral either — Greyhound Research sells analyst advisory and Kanerika sells AI implementation, and both are quoted arguing that governance work moves to the harness their kind of buyer pays for. No party in this story is positioned to say the numbers are wrong.
Firm on what was said, thin on what is true
What OpenAI announced — pricing, model name, channels, default-off enablement, Daybreak — is easy to stand behind; it is quoted directly and is the kind of fact a vendor cannot misstate for long. Everything about capability is a different matter, and the internal baseline conflict is a live reason for caution rather than a rounding note. One publisher, one lab's measurements, two analysts: enough to report confidently, not enough to score the model.