Build1 publisher3 min readPublished
Eleven tasks with pre-computed answer keys, three runs each, seven effort settings. Everything from low upward scored 33 of 33, so the only thing the top rung buys is the number on the launch page.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
`reasoning_effort` sets how much hidden thinking the model does before it answers, and that thinking is billed as reasoning tokens at the output rate [1]. So the rung you pick is a spend dial, not a mode switch, for every value except one.
The one exception is `none`, the only value that zeroes reasoning tokens [5]. It also fails 17 of the 33 verified runs [4]. Every other rung scored 33 of 33 [6]. The accuracy gap sits in that single step rather than spreading across the ladder, and that shape is what makes the pricing argument work: if `low` and `max` return the same answer on all 11 tasks [2], the extra tokens at the top are pure cost [3].
The 11 task answers were computed by exhaustive search on the author's own machine first, so the answer key cannot be wrong [7]. The tasks were things like a rule applied 40 times in a row, a counting problem under three constraints, a knapsack, a base conversion, and 7 to the power 222 modulo 1000 [8]. Those are deterministic, verifiable, short-horizon problems. For the ladder-buys-nothing conclusion to transfer to your workload, your tasks would have to be similar: single-shot, checkable, and already inside the model's competence at `low`. Long agent loops with tool calls are exactly the case this sweep does not cover, and they are the case OpenAI leads with, claiming Astra is ahead on agent-style work [11].
Then there is `disabled`, which is a rung the docs never mention and which the API accepts on Astra alone [9]. It spends 243 reasoning tokens against `low`'s 151 and costs 54% more per correct answer at the same accuracy [10]. Naming a setting `disabled` when it disables nothing is the kind of thing you only find by sending every value and reading the response codes.
The docs and the API disagree in both directions. The request-shape check advertises seven values including `none` and `minimal` [12], while the model page lists five and states the model does not support `none` [13]. `minimal` is advertised and rejected on every model [14]. `none` works despite the docs saying it does not [14].
The line that connects the benchmarks to the bill is OpenAI's own: "Evaluation scores are the maximum at any effort" [15]. The Artificial Analysis leaderboard lists each model once per effort setting, and Astra scores 54 at `xhigh` against 55 at `max` [16]. One point of aggregate score sits on top of a rung that, on these 11 tasks, changed nothing. On the same leaderboard's v4.2 version, Claude Fable 5.1 is first at 57 and Astra third at 55 [17], and OpenAI's own table puts Astra behind Fable 5.1, Opus 5 and Fable 5 on that row [18]. The Humanity's Last Exam row, where Astra trails Fable 5.1 by 7.8 points, appears in the table and not in the prose [19].
One more caveat on the security numbers: OpenAI says the cybersecurity scores were produced "without production safeguards", and the shipping model "will refuse" proof-of-concept exploit tasks [20]. That column reflects a configuration no buyer can actually run.
Cross-model, the same list prices put Astra at 1.57x GPT-5.6 Sol per correct answer at `low`, and 2.57x at `max`, on work both models get right [21]. Divide those and the effort ladder accounts for roughly 1.64x of the gap by itself [22]. The default worth defending is `low`, with `none` reserved for work where a wrong answer is cheap.
Ranked by verification strength, evidence, and original report placement.
reasoning_effort is the request parameter that sets how much hidden thinking the model does before it answers; that thinking is billed as reasoning tokens at the output rate.
On GPT-6 Astra, reasoning_effort max returns the same answer as low on every one of 11 verified tasks.
On GPT-6 Astra, reasoning_effort max costs 2.3x what low costs.
reasoning_effort none fails 17 of 33 runs (11 tasks, 3 runs each).
none is the only reasoning_effort value that zeroes reasoning tokens.
Every reasoning_effort rung other than none scored 33 of 33 verified runs.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Reproducible quotes, unreplicated runs
Two very different grades of proof sit in the same post. The quotable material stands on its own: the validation error listing all seven efforts is reproduced verbatim, the launch-page footnote about maximum-at-any-effort is a direct quote, and the $10/$50 against $4/$20 list prices are dated and attributed to OpenAI's model pages. The 33 runs behind the 2.3x arrive only as summary figures from a single author on a single machine, with the per-task table cited by reference and no logs published.
Shipping and priced, usage unseen
Astra and Sol are both live, listed at per-token prices on OpenAI's own pages as of 7 September 2026, and carried on the Artificial Analysis leaderboard with separate entries per effort, so the ladder is already in front of paying customers. Which rung anyone actually ships on stays invisible here — the only traffic described is the author's own 33 runs per model, and no usage disclosure or vendor telemetry appears.
Narrow sample, ladder-wide verdict
The measurement is narrower than the headline. Eleven puzzles with exact pre-computed answers is the workload least likely to separate the rungs, because a right answer is a fixed string and extra deliberation has nowhere useful to go. The post concedes the limit in one line, granting that a workload resembling the launch benchmarks may pay for max, and it tells readers to measure their own tasks at low first. The title prices the entire ladder off tasks designed to reward none of it, leaving the agentic strength OpenAI claims untested.
Anonymous auditor, motivated footnote
The author publishes under synthorai on dev.to and says nothing about what that account sells, so a reader cannot tell whether advocacy for the cheap rung also serves a commercial interest; the numbers arrive without a disclosure line. The vendor material under audit has an incentive that is easier to name: "Evaluation scores are the maximum at any effort" is precisely the footnote that lets a launch page quote the top rung, and the same page keeps the Humanity's Last Exam deficit in the table and out of the prose while flagging that its cybersecurity runs skipped production safeguards.
Arithmetic holds, sample is thin
We would repeat the documentation findings without hesitation: they are quoted strings, checkable in a minute, and internally consistent with the 200/400 results the post walks through. The money figures rest on thinner ground: they depend on one person's token accounting over 33 runs, and the two per-correct-answer multiples against Sol come from a passage that the text we hold cuts off mid-calculation. Treat them as a starting point for re-measuring your own workload rather than as a quotable price of reasoning.
science
Token efficiency absorbs GPT-6 Astra's 2.5x price increase inside the coding harness1 publisher
invest
Astra's 99.9% holds up only on the harness OpenAI ran itself1 publisher
leadership
The AI bill nobody reconciles: cost per finished task, not per million tokens1 publisher
build
Epoch's first-place ranking for GPT-6 Astra rests on a single coding score2 publishers
Publishers with included, body-backed reporting in this cluster.
1 article · September 7, 2026