Skip to content

Build1 publisher3 min readPublished

Luna lands at one tenth of Terra's price on both input and output tokens

OpenAI's GPT-5.6 cuts put Luna at $0.20 per million input tokens against Terra's $2, and the same factor of ten holds on output. Tier selection is now the largest single lever on a high-volume token bill.

The Engineer · Build desk

Illustration accompanying Luna lands at one tenth of Terra's price on both input and output tokens

What happened

  • The company says the new rates are beginning to roll out, including in AWS, and that Luna and Terra usage now consumes fewer credits than before.
  • Both models stay available across ChatGPT Work, Codex and the OpenAI API, so the cheaper rates apply to workloads already running there.
  • OpenAI says GPT-6 Astra usage is included within existing subscription allowances and that the model will also be reachable through the API.
  • Separately, the ChatGPT for Academic Researchers program announced free access for up to 100,000 researchers, starting with a lottery-selected cohort of 10,000.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Whoever owns the model-selection code now owns the largest line item in a high-volume token bill.
  • constraint Capturing the cheap tier's saving requires a validator good enough to distinguish Luna's failures from Terra's, because OpenAI has not published a rule for which option fits which task.
  • exposure Capacity plans written on the assumption that the research program's higher limits apply to ordinary accounts are exposed, since those limits ride on plan and eligibility.
  • precedent Rates that move per tier while model names stay fixed mean the routing table needs a re-audit after each announcement.

Terra lists at $2 per million input tokens and $12 per million output [2]. Luna lists at $0.20 and $1.20 [3]. Divide either pair and the factor is ten [15]. The symmetry is the useful part. Because the ratio holds in both directions, a task that moves from Terra to Luna sheds 90 percent of its token cost whatever the split between prompt and completion [16].

Astra breaks that pattern. Its standard rates are $10 per million input tokens and $50 per million output [8], five times Terra on input and about 4.2 times on output [18]. Routing away from Astra pays differently depending on whether the workload is prompt-heavy or generation-heavy. On a job that burns a million tokens each way, Terra costs $14, Luna $1.40, and Astra $60 [21].

The ten-to-one gap also decides how many retries are affordable. At equal token counts, ten Luna calls cost the same as one Terra call, so a pipeline that checks its own output can absorb nine bad Luna attempts before Terra is the cheaper route [17]. For that saving to reach a real invoice, two things have to hold: Luna's completions cannot run longer than Terra's on the same prompt, and whatever grades them has to be cheap enough not to eat the difference.

Fast mode is sold at a published multiple of the base price. On Sol it is quoted at up to 2.5 times faster for twice the price, and priority requests use that mode automatically [6]. The switch happens for you; priority traffic buys the double rate on your behalf. Astra's Fast mode is described as roughly twice as fast for twice the price [8], so on the quoted figures only Sol's option returns more speed than money [20]. Sol's base rate appears in neither dev.to report, so the surcharge cannot be put in dollars from them [23].

The percentages and the list prices reconcile if each cut applied equally to input and output, which puts Luna at $1.00 and $6.00 before the change and Terra at $2.50 and $15.00 [19]. Both dev.to posts attribute the figures to OpenAI's official price-performance announcement [22]. The same reporting says OpenAI has not defined which GPT-5.6 option is best for a given task or published a universal formula for model selection [11], and that nothing requires moving existing Luna or Terra workloads because of the price update [12].

dev.to frames the changes as consistent with public remarks from OpenAI chief executive Sam Altman about offering the strongest intelligence and price combination at each point along a Pareto-optimal frontier [13], which the same piece defines as choices where improving one dimension such as cost means giving up something else such as speed or capability [14]. The capacity news comes with a caveat. dev.to cautions that the higher limits described in the academic researchers program are tied to particular plans and eligibility contexts, and stop short of confirming that every API or subscription customer gets them [10].

What to watch

  • A published base rate for Sol, which would let the Fast mode surcharge be stated in dollars per million tokens.
  • Whether the AWS rollout lands the same list rates as the direct OpenAI API, and on what date each region gets them.
  • OpenAI guidance or an eval that maps where Luna's output quality falls short of Terra's on specific task types.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories