Build2 publishersIndependently confirmed3 min readPublished Updated
GPT-6.1 Sol matched GPT-6 Astra across 15 test runs at 18% of the cost
OpenAI's GPT-6.1 Sol matched GPT-6 Astra on all 15 runs of a New Stack test, costing $2.66 against Astra's $14.77. Both models scored perfectly, so the run confirms the price gap on real token counts while leaving the accuracy ceiling untested.
The Engineer · Build desk

What happened
- OpenAI released GPT-6.1 Sol on September 29, a week after GPT-6 Sol, and markets it as a cheaper near-match for GPT-6 Astra.
- The three tests were CI triage across 40 failed job logs, seven postmortem questions on 3,664 log lines, and a dependency resolver graded by 120 hidden tests.
- Astra was faster on CI triage, averaging 24 seconds a run to Sol's 31, while Sol was faster on the two longer tests.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Because the rate card fixes every per-token ratio, a team can predict its own Sol-versus-Astra bill from output-token counts on a small side-by-side run of its own prompts.
- constraint A perfect tie sets no upper bound on either model, so a team with tasks harder than these still has to find for itself where Sol starts to fall behind.
- cost The one-tenth cached ratio lowers Sol's relative cost only where cached input stays most of the bill; output-heavy jobs stay near one-fifth of Astra's price.
OpenAI's pitch is "near-Astra intelligence for a fifth of the price" [4]. The price part of that line comes straight from the rate card. Sol's input and output rates are each exactly one-fifth of Astra's [6]. For the same token counts on uncached traffic, Sol bills 20% of what Astra does [23]. The prompts were identical, so the ratio could only move when the two models wrote different amounts of output.
On CI triage and the incident logs, the two models' output counts were close [10][12]. The computed cost ratios came to 19.5% and 19.9% [17][18]. The resolver spec pulled the ratio down to 15.6% [19]. On that test Astra wrote 25,207 output tokens a run to Sol's 19,637 [13], and output was about 98% of Sol's resolver bill [22]. Take the resolver runs out and the 15-run ratio rises from 18% to about 19.8% [7][20].
The reported costs price every input token at the full rate. On the incident logs, 113,966 input tokens at $2 per million plus 8,316 output tokens at $10 per million comes to $0.311, against the reported $0.31 [12][18]. OpenAI halved Sol's cached-input price to $0.10 from GPT-6 Sol's $0.20 [5]. That is one-tenth of Astra's $1 cached rate [6]. With the whole prompt cached, the same incident-log run would cost about $0.095 on Sol and $0.541 on Astra [21]. The ratio would be 17.5%. Once input is cheap, output makes up most of both bills, and output is still priced at one-fifth [21][6].
Both models were perfect on all 15 runs [7]. Astra and Sol are both OpenAI models, so either result in a tie suits the vendor [4]. A tie at 100% shows Sol can handle these three tasks at maximum effort. The tasks are hard enough to catch mistakes. In the reviewer's earlier round, GPT-6 Sol missed a customer on two incident-log runs, miscounted failed checkouts on another and crashed one resolver run with a stray parenthesis [16]. "My tests can't show the two models are equal on everything," the reviewer wrote [14].
For the 18% to hold on another team's work, that work has to look like these runs. They used identical prompts, maximum reasoning effort and a 64,000-token output cap [8]. OpenAI's claim of fewer factual errors is made at low reasoning effort [3]. A team trying to cut costs further would probably try that setting next. The resolver was written without running any code [9]. OpenAI's near-match claim also covers computer use [2].
"OpenAI's claim held up in my tests," the reviewer wrote [15]. For work shaped like these three tasks, I think Sol is the right default. I'd keep Astra for whatever a team's own eval shows Sol getting wrong. Without that eval, the evidence is 15 runs on one reviewer's prompts [7].
What to watch
- A Sol-versus-Astra comparison at low reasoning effort, the setting where OpenAI makes its factual-error claim, with output tokens reported for both models.
- Head-to-head results on computer use, the part of OpenAI's near-match claim outside these three tests.
- Tests on tasks hard enough that at least one model fails some runs, so an accuracy gap between Sol and Astra can be measured.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+5
- Incentives35
- Confidence55
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
OpenAI launched GPT-6.1 Sol on September 29, a week after it launched GPT-6 Sol, and marketed 6.1 as an upgrade to 6 and as a cheaper near-match for Astra.
ReportedSupportedSource: The New Stack2 sources— create a free account to open themView cited source - [2]
OpenAI says GPT-6.1 Sol nearly matches Astra on agentic coding, computer use, and professional work.
ReportedSupportedSource: OpenAI, via The New Stack2 sources— create a free account to open themView cited source - [3]
OpenAI says GPT-6.1 Sol makes fewer factual errors than GPT-6 Sol at low reasoning effort.
ReportedSupportedSource: OpenAI, via The New Stack2 sources— create a free account to open themView cited source - [4]
OpenAI calls GPT-6.1 Sol "near-Astra intelligence for a fifth of the price."
ReportedSupportedSource: OpenAI marketing, via The New Stack2 sources— create a free account to open themView cited source - [5]
OpenAI halved cached-input pricing, from GPT-6 Sol's $0.20 to $0.10 per million tokens.
ReportedSupportedSource: The New Stack2 sources— create a free account to open themView cited source - [6]
GPT-6.1 Sol costs $2 per million input tokens, $10 per million output tokens and $0.10 per million cached input tokens; GPT-6 Astra costs $10 per million input tokens, $50 per million output tokens and $1 per million cached input tokens.
ReportedSupportedSource: The New Stack2 sources— create a free account to open themView cited source - [7]
Both models were perfect on all 15 runs; GPT-6.1 Sol's runs took 49 minutes 54 seconds and cost $2.66 in total, Astra's took 1 hour 6 minutes 55 seconds and cost $14.77; Sol cost 18% of what Astra did.
ReportedSupportedSource: The New Stack reviewer2 sources— create a free account to open themView cited source - [8]
The reviewer called both models through the OpenAI Responses API with identical prompts, reasoning effort set to max, and a 64,000-token output limit.
- [9]
The three tests: CI triage (40 failed CI job logs plus a runbook with conditional rules; retry, block or page for each); incident logs (3,664 lines from five services during a two-hour outage, seven postmortem questions); resolver spec (write a dependency resolver for a fictional package manager from a two-page spec, without running any code, graded by 120 hidden tests).
- [10]
On CI triage both models made all 40 calls correctly on all five runs; both read 3,867 input tokens per run; output was 1,381 tokens for GPT-6.1 Sol and 1,431 for Astra; Sol cost about 2 cents per run and Astra 11 cents.
- [11]
Astra was faster on CI triage, averaging 24 seconds per run to GPT-6.1 Sol's 31 seconds; GPT-6.1 Sol was faster on the two longer tests.
- [12]
On incident logs both models answered all seven questions correctly on every run; each run read 113,966 input tokens; output was 8,316 tokens for GPT-6.1 Sol and 8,539 for Astra; Sol cost $0.31 per run and Astra $1.57; Sol averaged 2:20 per run, about 19% faster than Astra's 2:53.
- [13]
On the resolver spec both models passed all 120 hidden tests on every run; Astra wrote 25,207 output tokens per run to GPT-6.1 Sol's 19,637, or 28% more; Astra cost $1.28 per run and Sol $0.20; Sol averaged 7:07 per run, about 30% faster than Astra's 10:06.
- [14]
"My tests can't show the two models are equal on everything"
- [15]
"OpenAI's claim held up in my tests."
- [16]
In the reviewer's earlier tests, GPT-6 Sol missed a customer on two incident-log runs, miscounted failed checkouts on another, and left a stray parenthesis in one resolver run that crashed the resolver.
- [17]
Computed CI triage cost per run: Sol about $0.0215, Astra about $0.1102, a ratio of about 19.5%.
- [18]
Computed incident-log cost per run at uncached list prices: Sol about $0.311, Astra about $1.567, a ratio of about 19.9%; these match the reported $0.31 and $1.57, so the reported costs price all input at the uncached rate.
- [19]
Resolver spec cost ratio, Sol to Astra, is about 15.6%.
- [20]
Excluding the resolver runs, Sol's total cost would be about 19.8% of Astra's.
- [21]
If the whole incident-log prompt were billed at cached rates, a run would cost about $0.095 on Sol and $0.541 on Astra, a ratio of about 17.5%, with output making up most of each bill.
- [22]
Output tokens account for about 98% of Sol's $0.20 resolver run cost.
- [23]
Because Sol's input and output rates are each one-fifth of Astra's, for identical uncached token counts Sol's bill is 20% of Astra's; any other ratio comes from different token counts or caching.
Sources
2 independent publishers whose own reporting we read for this story.
- dev.toGPT-6.1 Sol vs GPT-6 Astra: One-Fifth Price Math for Agents
1 article · October 8, 2026
- thenewstack.ioGPT-6.1 Sol vs. GPT-6 Astra: Same accuracy at 18% of the cost
1 article · October 7, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.