Invest1 distinct publisher3 min readPublished
Vendor benchmark tables are dated snapshots, and the GPT-6 Astra launch shows how much can move inside one afternoon without the headline score changing, which matters for anyone scoring a purchase off one.
The Investor · Invest desk

Compiled by The InvestorSomething wrong?How this is made
The distance between two numbers is the thing a procurement table exists to supply, and that distance moved while neither model did. In the 2:23 p.m. Internet Archive snapshot of the launch post, Astra led Anthropic's Fable 5.1 on FrontierMath Tier 4 (v2) by 9.8 points, 97.6 against 87.8 [9][10][1]. In the 5:17 p.m. snapshot it led by 19.6, 97.6 against 78, exactly twice the original margin, with Astra's own figure untouched throughout [9][10][2]. That shift happened in two hours and fifty-four minutes [8]. The margin now reads 14.6 [3].
An OpenAI spokesperson told Fortune that most evaluations carry noise within a few percentage points depending on the exact checkpoint, scaffold and evaluation run, and that the launch-blog fixes were made so users could make meaningful comparisons [5]. A competitor's score falling 9.8 points and then recovering 5 of them sits well outside a band described as a few [7][9].
The same afternoon halved Astra's reported hallucination rate from 4.2% to 2% alongside four other changed metrics, and cut GPT-5.6 Sol's from 12.2% to 9.4%, before both returned to where they started [6][7]. Read as a ratio, which is how anyone comparing generations reads it, Astra went from 2.9 times cleaner than its predecessor to 4.7 times cleaner and back [4]. On ARC-AGI-3 the embargoed draft sent to media said 98.6% and the live page says 99.99% [12], a 1.4-point move that is comfortably inside OpenAI's stated noise band and simultaneously a 140-fold reduction in residual error [5]. Near the ceiling, a point is not worth what a point is worth in the middle.
The pattern is not uniformly self-serving. Sol's internal ExploitBench score was raised from 5.5% to 11.5%, a doubling that flatters the predecessor rather than the new model [8][6], and Sol's hallucination figure improved in the same pass [7]. More to the point, the flattering numbers did not stick: the hallucination rates are back at 4.2% and 12.2%, and Fable is back at 83% [6][7][10]. A lab willing to walk a number back this way is doing something closer to error correction than a lab that lets an inflated figure ride.
Which leaves the ExploitBench disclosure as the load-bearing item, or rather the more useful version of it: OpenAI says it is investigating reverting Sol's 11.5% because that result reflects a reasoning level not commercially available for Sol [8]. A published figure describing a configuration nobody can buy is an accurate measurement of something else entirely, not a noisy measurement of the product, and no error bar reconciles the two.
What would falsify the sceptical read is disclosure rather than restraint. If each cell arrives with a checkpoint, a scaffold, a reasoning level and a run date, and independent reruns land inside the few-point band OpenAI describes [5], the table becomes usable with a tolerance attached. Absent that, the observed movement of 9.8 points on a comparator [7] exceeds most of the margins a buyer would act on, which makes the vendor table a shortlist filter and leaves the scoring to a harness the buyer pays for.
Ranked by verification strength, evidence, and original report placement.
OpenAI has changed several evaluation benchmarks for its GPT-6 Astra model since first publishing a blog post announcement mid-afternoon on Sept. 3; in some cases the updated numbers showed Astra performing better while numbers for models from Anthropic got worse.
OpenAI originally planned for the blog post to go live at 2 p.m. ET, but it took almost another two hours before it was widely viewable online.
OpenAI published the blog shortly after 2 p.m. but retracted it for reasons the company said it could not disclose, which it said were unrelated to the benchmark performance figures; it first told Fortune it was a bug in the content management system, then an internet outage. On republishing, the post had different evaluation metrics that seemed to favor Astra.
An OpenAI spokesperson told Fortune: "We care deeply about getting evaluations right. Most evaluations have noise within a few percentage points based on the exact checkpoint, scaffold, and evaluation run used in reporting. For our launch blog, we made fixes to ensure the numbers represent our best estimate of available model performance, so that users can make meaningful comparisons."
Astra's reported hallucination rate was 4.2% in the first Internet Archive snapshot at 2:23 p.m. and remained so through a fifth snapshot at 3:11 p.m. ET; in the sixth snapshot, taken at 5:20 p.m., it was halved to 2%, along with four other changed metrics. As of Fortune's writing the rate was back up to 4.2%.
The hallucination score for Astra's predecessor GPT-5.6 Sol went down from 12.2% to 9.4% in the same change, and as of Fortune's writing was back up to 12.2%.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 4, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
invest
Astra's 99.9% holds up only on the harness OpenAI ran itself1 distinct publisher
product
OpenAI ships a computer-use agent it classifies as a critical cybersecurity capability7 distinct publishers
invest
OpenAI rates GPT-6 Astra capable of hacking hardened systems without human guidance1 distinct publisher
invest
Sanders and Casar attach a 20-year prison term to building superintelligence1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Dated captures, one newsroom
The spine of this story is documentary: six Internet Archive captures of the same page with times attached, an embargoed draft in Fortune's possession, and two on-record OpenAI statements. Anyone can pull the same captures and check the numbers, which is stronger than most single-publisher reporting gets. What is missing is a second party measuring the models, so every score in play is still OpenAI's own.
Shipped; uptake unreported
Two things here count as real-world signal: Astra launched on Sept. 3, and the Arc Prize Foundation ran it independently on its own benchmark. Beyond that the reporting carries no deployments, pricing, customer counts or usage figures, and the score changes describe a web page rather than anything running in production. The low number reflects what was reported, not evidence that the model is going unused.
Table outruns its own caveat
OpenAI puts evaluation noise at a few percentage points, then cut a rival's FrontierMath score by 9.8 points and halved its own hallucination rate before restoring both. Astra's ARC-AGI-3 figure went to 99.99% in public while the benchmark's author measured 63% under a standard harness. Fortune's writing stays inside what the captures show, so the overstatement being scored belongs to the vendor's table rather than to the reporting on it.
Launch-day comparison table
The page being edited is marketing for a model competing directly with Anthropic, and the edits went in the flattering direction on the metric the post leads with. OpenAI's own process puts different research teams in charge of different metrics before a central team publishes, and one revision used a reasoning level Sol does not sell. Fortune has an interest here too: it held the embargoed draft, which is what made the discrepancy story possible.
Checkable, still moving
The timeline and the quotes are solid enough to hold up to a reader repeating the archive lookups. Confidence is capped by three gaps: no second newsroom, no response from Anthropic on a score that dropped nearly ten points, and figures that were still changing as Fortune filed, with OpenAI saying another revert is under consideration. Any assessment of what the current table says has a short shelf life.