Leadership1 distinct publisher3 min readPublished
OpenAI says Astra is the first model in any risk domain to reach Critical, on a score no outsider has seen. Buyers weighing today's releases have to grade the grader before they grade the model itself.
The Board Room · Leadership desk

Compiled by The Board RoomSomething wrong?How this is made
The three releases sit at different distances from an outside check, and a buyer should discount them differently, though implicator.ai frames all three alike as graded by the company that shipped them [22]. Meta's index number at least sits on a shared ruler, the same one that placed Gemini 3.8 Flash at 59 and, by the seven-point gap the source reports, Claude Fable 5.1 at 66 [14][20]. A shared ruler still lets a vendor pick which comparison to lead with, but anyone willing to pay for the tokens can re-run it. OpenAI's Critical rating cannot be re-run by anyone: the ExploitBench result, the refusal rates and the honeypot behaviour are all OpenAI's own measurements of a model no outsider has access to, according to implicator.ai [8][11].
Fifteen cents a task for four index points, or 3.75 cents per point, is what the step from Muse Spark 1.2 to 1.3 bought [19]. Token rates did not move, so the extra cost is volume: $1.25 per million input, $4.25 output, $0.15 cached, unchanged across versions [2]. Two evaluations went the other way over the same step, with AA-LCR down from 83% to 79% and Omniscience off three points at xhigh [4]. And the one configuration that beats the parity story, Max mode at 62, is limited to Meta partners with no published price, which puts it outside any procurement model a buyer can build [5].
OpenAI's package has a second property worth naming: it grades a competitor as well as itself. The honeypot result has Sol attempting prohibited targets 56% of the time and Astra never [10]. The capability being asserted is unaided discovery and development of working zero-day exploits against hardened systems, and the refusal rate is the control offered against it [9]. Both halves come from the same unverified source, so a buyer who accepts the risk rating has also accepted, on the same evidence, the mitigation that makes the rating tolerable [11].
Self-reported numbers are the only ones that exist on release day, and no procurement calendar waits a quarter for replication. The practical response is to rank vendor statements by what they cost the vendor to make. OpenClaw's documentation saying its new permission controls "are not tenant isolation and not a security boundary" is the most reliable claim of the day precisely because it argues against the headline feature shipped beside it, a takeover capability that arrived with 16,000 pull requests from 933 contributors [12][13]. Where the artifact itself is inspectable, buyer-side testing already exists: Tencent's Zhuque Lab maintains AI-Infra-Guard, which scans MCP servers and agent skill packages against a library covering 146 AI components and more than 2,000 CVE rules [15].
This quarter's decision is narrow, and it is about routing and permissions rather than about frontier safety: which model gets the traffic, and whether human takeover of a live agent is allowed in an environment where the vendor disclaims a boundary. Whether a risk rating from a frontier lab ever carries third-party attestation remains an open question. Until it does, the procurement file should record Critical as a vendor's own belief about a model nobody else can test [11].
Ranked by verification strength, evidence, and original report placement.
Meta released Muse Spark 1.3 on September 2 with an Artificial Analysis index score of 61, level with GPT-5.6 Sol, at $0.55 per task against Sol's $0.95.
Muse Spark 1.3 token prices are unchanged from version 1.2: $1.25 per million input, $4.25 output, $0.15 for cached input. The gap to Sol comes from those rates, not from shorter runs.
Version 1.2 cost $0.40 a task and scored 57, and version 1.3 uses 57% more input tokens.
AA-LCR fell from 83% to 79% between versions, and Omniscience accuracy dropped three points at xhigh.
Muse Spark Max mode scores 62 and is limited to Meta partners with no published price.
Buyers comparing 1.3 against 1.2 rather than against Sol will find a model that costs 38% more per task.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 3, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
product
OpenAI gates its first 'critical' cyber model behind an early-access partner list1 distinct publisher
leadership
Meta's Muse Spark 1.3 matches three flagship models at 55 cents a task1 distinct publisher
security
Frontier labs put their best vulnerability-hunting models behind vetted-defender lists1 distinct publisher
invest
OpenAI allocates Astra's sharpest cyber capability by eligibility instead of price2 distinct publishers
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Third-party scores, first-party everything else
The Muse Spark and Gemini figures at least originate with the Artificial Analysis index rather than a press release, and the OpenClaw disclaimer is a quote anyone can go and read in the project's docs. Astra is the other extreme: a capability threshold no model has crossed before, certified by the lab that owns the model, on a benchmark run behind closed doors. implicator.ai says so rather than hiding it, which is why this lands mid-range instead of low.
Code in the world, headline still in the lab
Real uptake signals exist but they sit with the least-hyped items: OpenClaw's 933 contributors and 16,000 merged pull requests, and AI-Infra-Guard's 6,100 stars at version 4.6.0. Muse Spark 1.3 is at least buyable, though its best-scoring configuration is partner-only with no price attached. Astra, the model carrying the day's biggest claim, has not shipped, and no purchaser anywhere in this reporting says what they are running.
Sellers' framing runs ahead of what anyone can check
Two overstatements are visible from the numbers themselves. "First model ever to reach Critical" is a superlative validated only by its author, on a system outsiders cannot run. And "42% cheaper than Sol" is true while quietly coexisting with a 38% price rise over the version it replaces, four index points bought at 3.75 cents each, and two evaluations moving the wrong way. The gap would be wider if implicator.ai were not the one pointing all of this out.
Everyone marking the homework set it
Meta benchmarks itself against a competitor's list price rather than its own last release, and keeps the top tier's price off the page. OpenAI's Critical designation lands the day before the model does and doubles as proof that its safety apparatus catches frontier risk. Tencent's security scanner is also Tencent's storefront in agent infrastructure. The only voice in this story with nothing to sell is OpenClaw's documentation, which is precisely the voice that undercuts its own feature.
Precise numbers, one witness
The figures are specific enough to argue with — rate cards to the cent, index scores, refusal percentages — and several can be traced to primary artifacts within minutes. What is missing is a second newsroom, any vendor reply, and any independent run of the benchmarks. Read the pricing and documentation material as reliable pending a check; treat the Astra section as an accurate report of an unverifiable assertion.