Science2 publishers3 min readPublished
OpenAI measures its 50% GPT-6 price cut against the previous generation's promotional rate
Sol lists at $2 and $10 per million tokens and Luna at $0.10 and $0.50. The per-task savings OpenAI published come mostly from the lower price, and the cheaper tier scores below its predecessor on computer use.
The Scientist · Science desk

What happened
- OpenAI added two cheaper tiers to the GPT-6 family, Sol and Luna, and said it cut their API prices by half compared with what it describes as GPT-5.6 promotional pricing.
- Sol lists at $2 per million input tokens and $10 per million output and Luna at $0.10 and $0.50, one fifth and one hundredth of Astra's $10 and $50 respectively.
- xAI priced Grok 4.7 at $2 and $6 on September 21, Anthropic released Opus 5.5 at $4 and $20 the next day, and OpenAI published Sol and Luna ninety minutes after that.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- cost Anyone who budgeted this quarter on GPT-5.6 promotional rates gets half the per-token price and an unquantified change in tokens per task, so the saving that shows up on the invoice is not 50 percent.
- constraint A team running desktop agents on GPT-5.6 Sol cannot treat the cheaper tier as a drop-in, because the same workload gives up 5.2 points of measured completion for the lower price.
- decision OpenAI's own reported drop to 1.3 percent coding deception in a tier priced at a fifth of Astra makes it harder for any vendor to argue that frontier-tier prices are what buys behavioural safety.
Per-token price and cost per task are different numbers, and the second one lands on the invoice. OpenAI's own figures show the gap. On AutomationBench, GPT-6 Luna at high effort scored 5.4 percentage points above its predecessor at 58 percent lower cost per task [15]. The list price fell by half, against GPT-5.6 promotional pricing in OpenAI's wording, and the post does not give the standard rates [2]. Cost per task is price times tokens consumed. If both published figures hold and the mix of input and output tokens is unchanged, Luna spent about 16 percent fewer tokens per task than GPT-5.6 Luna did: 0.42 divided by 0.50 is 0.84 [16].
The effort setting accounts for the rest. Every comparison OpenAI leads with is run at xhigh or max effort [20][22]. Forkast puts Sol at 33.2 percent on AutomationBench at $0.27 per task, against Claude Opus 5's 26.9 percent at roughly $3 [17]. That matches OpenAI's claim of 9 percent of Opus 5's cost per task [18]. At Sol's $10 per million output tokens, $0.27 covers at most 27,000 output tokens, so one "task" on that test is a long run [19]. OpenAI also reports Luna at max effort scoring 66.6 percent on DeepSWE v1.1, 93 percent cheaper per task than Opus 5 and 96 percent cheaper than Fable 5 [21].
On output tokens, rivals undercut OpenAI. Sol's $2 input price matches Grok 4.7 [8], but Grok's output price is $6 and Sol's is $10, so Sol costs 67 percent more per output token [12]. Luna's $0.10 input undercuts Xiaomi's MiMo-V2.6 Flash at $0.14, while Luna's $0.50 output is 79 percent above MiMo's $0.28 [13]. Forkast writes that Luna "sits below every frontier-class model on the market" [11]. For output-heavy work, its own price list says otherwise.
On computer use the cheaper model did worse. Forkast reports Sol at xhigh scoring 60.5 percent on OSWorld 2.0 against GPT-5.6 Sol's 65.7 percent [23], a fall of 5.2 points [24]. OpenAI says Astra remains its best model for computer use and that Sol and Luna "offer more cost-efficient performance than their predecessors" [25]. Both can be true at once, and a team pointing an agent at a desktop is trading measured completion for the lower price.
OpenAI's internal factuality evaluation is built from de-identified real conversations in which users flagged mistakes, and on it Sol makes about half as many mistakes as its predecessor [26]. That set is selected for having gone wrong, so halving the error count on it is not the same as halving errors across all traffic. OpenAI also reports Sol's coding deception rate at 1.3 percent, down from 10.4 percent, and broken tool disclosure failures falling from 77.5 percent to 4.9 percent [28]. Those are its own evaluations.
Forkast frames the week as the end of a price truce: xAI shipped Grok 4.7 on September 21 at $2 and $6, nine days after Elon Musk endorsed Dario Amodei's call for a coordinated slowdown [9], Anthropic answered with Opus 5.5 at $4 and $20, and OpenAI published ninety minutes later [10]. An antitrust suit filed September 18 alleged the slowdown was an output-restricting cartel, and Forkast argues the price war that followed "makes that argument harder to dismiss" [29]. The order of the releases does not tell you what any of these models cost to serve. OpenAI's stated reason is narrower: improvements in caching and inference "let us serve these models at lower cost" [3].
OpenAI's post also puts a figure on sustained agent use inside the company. Valued at API prices, daily token usage exceeded $600 for the median researcher and $7,000 for researchers at the 90th percentile, and the company says its internal usage "has grown exponentially" [30].
What to watch
- Independent reproduction of the OSWorld 2.0 and DeepSWE v1.1 figures by anyone other than OpenAI.
- Whether OpenAI publishes GPT-5.6 standard rates, so the 50 percent can be checked against a non-promotional baseline.
- Whether the September price moves are cited in the antitrust case filed on September 18.