Product1 distinct publisher3 min readPublished
The company says its next major model clears its own Critical cybersecurity threshold, so the strongest security functions ship to selected partners while everyone else reads the capability in a blog post.
The Product Desk · Product desk

product
OpenAI prices its own guardrails: 20% more compute, plus a two-week training pause1 distinct publisher
invest
Bailey's letter to the G20 warns AI-driven leverage and market concentration could amplify a crash7 distinct publishers
invest
The AI trade's weak link is the buyer: Anthropic's best model took 6% of its tokens1 distinct publisher
build
Grok 4.6 lands in Copilot two days after launch, and the model picker becomes a procurement problem1 distinct publisher
Compiled by The Product DeskSomething wrong?How this is made
A security engineer who reads the Sept 1 post and then opens their API console will not find the capability described in it, because the announced plan sends the most advanced cyber functions to a closed group of select testing partners first, and expands access for defensive use "eventually" [4].
The tier definition is worth reading before deciding how much of this is packaging. A model reaches Critical cybersecurity, OpenAI writes, if it can "identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention," or can "devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal" [6]. That definition describes something closer to a staffing plan than a feature, and it is the reason a gate exists at all.
The calendar matters as much as the definition. Aug 7 was the hedge: new security controls, some internal development work paused, and an admission that the company could not rule out critical cyber capability [2]. Sept 1 was the confirmation [3]. Twenty-five days between "cannot rule out" and "meets the threshold" [7]. Any team whose model-approval process runs on a quarterly cycle should hold it up against that number.
The usual pattern, a published capability landing in the API within a release or two at a listed price with the buyer deciding whether to pay, does not apply here. Instead, the capability is documented publicly and provisioned to a partner list of unstated size, with no route in that OpenAI has published; the closest hint is its earlier commitment to work with "relevant government agencies and select AI safety organizations" on testing this model [8]. Mashable reports the company declined to answer its questions about Astra [9].
The gate is not only about who might abuse the thing. OpenAI names two safeguard objectives: keeping malicious actors from using the model, and stopping Astra from taking "unauthorized, misaligned actions" [11]. The second one holds no matter how clean the requester is, and it is what keeps the tightened sandboxes and the chain-of-thought interruption in place [10]. A defensive-security buyer with impeccable paperwork is still on the wrong side of objective two.
The pressure this is meant to answer is already visible elsewhere. Mashable reports that some zero-day bug bounty programmes have shut down under the volume of AI-discovered bugs [14]. Naming is unresolved too: nobody outside the company knows whether this ships as GPT-5.7, as the start of GPT-6, or under some other label [12], and enterprise entitlements are written against model names, not against blog posts.
Security capabilities on a roadmap tend to fall into one of four categories: provisioned to your account, provisioned to someone else's account, described in a vendor post, or inferred from a benchmark. Only the first category holds a commitment a team can defend. By OpenAI's own account, Astra's cyber capability sits in the second [4], which means any line item reading "the model will triage this for us" carries a dependency the vendor has already said is not yet available.
Ranked by verification strength, evidence, and original report placement.
On Aug. 1, OpenAI confirmed the existence of a model called Astra, describing it as "our next major model," after revealing that an internal model of that name had solved 10 major open math problems, some unresolved for decades.
On Aug. 7, OpenAI announced that Astra had developed cyber capabilities advanced enough that new security controls were necessary and some internal development work would be paused, saying it could not "rule out critical cyber capabilities under our Preparedness Framework."
On Sept. 1, OpenAI confirmed in a blog post that Astra meets the "critical" threshold for cyber capabilities and that the model will be "available soon."
For safety reasons OpenAI plans to restrict Astra's most advanced cybersecurity capabilities to a closed group of select testing partners at launch, with access expanded later so the model can be used for defensive purposes.
OpenAI previously deemed GPT-5.6-Sol a "high" cybersecurity risk, making Astra the first OpenAI model classified as a genuine "critical" risk.
OpenAI's stated definition: a model reaches the Critical cybersecurity threshold "if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal."
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 1, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Well quoted, single voice
Follow any fact in this story back and it ends in the same place: an OpenAI blog post, relayed by Mashable. The Critical definition is reproduced verbatim, which is worth something. But no outsider has run Astra against that definition, not one testing partner or agency is named, and the company won't answer Mashable's questions. Accurate transcription of a single interested party is not corroboration.
Nothing shipped, sharp end fenced off
The only concrete deployment fact available is a restriction. Astra is unreleased, its date is the phrase "available soon," and its strongest security functions go to a closed list of unnamed testers while everyone else reads the capability claim. August's news was work being paused, not work going out.
Self-graded danger, pre-product
"Critical" is OpenAI's word, measured against OpenAI's rubric, applied to a model nobody outside has run — and the danger label arrives before the product does. Mashable deserves credit for saying out loud that when Anthropic declared Mythos too dangerous to release the warning doubled as marketing, then applying only half of that skepticism here. The capability may be entirely real; what is overstated is how much of it has been established rather than announced.
Grader, gatekeeper and vendor are one party
The same company writes the threshold, decides its model clears it, and chooses who may use the capability that clearing it implies — then sells the model. Locking the sharpest functions to select partners is defensible safety practice and manufactured scarcity in the same gesture, and declining to answer questions keeps the whole frame inside the blog post. The one countervailing incentive is real too: a company claiming critical cyber capability invites regulators to take it at its word.
Firm on the paper trail, blank on the substance
Dates, quotes and the classification history are solid and easy to check against the posts themselves. Everything past that is thin: one publisher relaying one issuer, and the question that actually matters — whether Astra really writes working zero-days against hardened targets unaided — is the exact thing no evidence in this reporting touches.