Skip to content

Invest1 publisher3 min readPublished

OpenAI cuts the input price of its Luna model line by 90% in 54 days

OpenAI has cut input prices on its Luna models from $1.00 to $0.10 per million tokens since July 30, over two rounds of reductions. Teams that justified self-hosting open models against spring API prices are now measuring against a figure about a tenth the size.

The Investor · Invest desk

Illustration accompanying OpenAI cuts the input price of its Luna model line by 90% in 54 days

What happened

  • Anthropic made Claude Sonnet 5's introductory price permanent on August 10, at $2 per million input tokens and $10 per million output tokens.
  • Anthropic answered with Claude Opus 5.5 at a list price of $4 input and $20 output, with efficiency gains the report says save users around 40%.
  • Open-source models still hold significant volume share, particularly on OpenRouter, where developers mix and match models.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • decision A self-hosting budget benchmarked against July's $1.00/$6.00 Luna pricing now faces a rival 90% cheaper on input and about 92% cheaper on output, so the build case has to be redone with current prices.
  • contradiction The report's claim that Chinese pricing undercuts US models by a wide margin holds against GPT-6 Sol but fails on input against GPT-6 Luna, where DeepSeek charges 40% more.
  • exposure Buyers moving routine work onto cheap proprietary tiers depend on prices set while both labs trade per-call margin for usage ahead of expected IPOs, and nothing stops those prices from rising later.

Working backwards from the July cut gives the starting price. An 80% reduction that left GPT-5.6 Luna at $0.20 per million input tokens and $1.20 per million output tokens means it cost $1.00 and $6.00 the day before [1][12]. GPT-6 Luna arrived on September 22 at $0.10 and $0.50 [4]. So over 54 days the input price of OpenAI's Luna line fell 90% and the output price fell about 92% [13][14]. Terra, the higher-end model, went from $2.50 and $15 to $2 and $12 [2][15].

The mid-tier has settled on a single price. Claude Sonnet 5 and GPT-6 Sol both charge $2 in and $10 out [17]. Anthropic lists its more expensive Claude Opus 5.5 at $4 and $20, and CryptoBriefing says efficiency gains cut the effective bill by around 40% [5].

CryptoBriefing argues the cuts are steep enough to make developers ask whether self-hosting open models is still worth the trouble [11]. The report includes no GPU, power or staffing costs, so it cannot settle a breakeven. It does fix the other side of the comparison. A self-hosted model justified against July's $1.00 and $6.00 Luna is now up against $0.10 and $0.50 [12][4]. That competition sits in the cheap tier. According to the report, enterprises now send summarisation, classification and basic customer interactions to cheap proprietary models and keep premium models for specialised work [6].

The two American labs are not the only source of cheap tokens. DeepSeek V4 Flash costs roughly $0.14 and $0.28, a price CryptoBriefing says undercuts American counterparts by a wide margin [8]. Against Sol that holds: Sol charges about 14 times as much for input and 36 times as much for output [19]. Against GPT-6 Luna it fails on input, where DeepSeek charges 40% more, and only DeepSeek's output is cheaper, by 44% [16]. The report adds that some Chinese models approach American mid-tier capability at five to nine times lower cost [9].

Usage is the other test of the thesis. Open-source models still command significant volume share, particularly on OpenRouter, according to the report [7]. If that share holds through the months after September 22, developers are picking open models for reasons that per-token price does not capture.

The sellers' motive is the main risk to the buy case. Both labs are widely expected to pursue IPOs, and the report describes the cuts as a bet that higher volume will eventually make up for thinner margins on each call [10]. OpenAI and Anthropic are giving up margin this year to buy usage [10]. A team that retires its own inference hardware on these prices also gives up the fallback it would need if they rise after a listing.

In my view the per-token case for self-hosting routine workloads is weaker than it was in July. A build case still quoting spring API prices is off by a factor of ten on input [13]. The counter-case is the IPO: prices set to lift adoption before a listing are the least durable number in the comparison [10]. I would be wrong if open-source volume on OpenRouter holds [7], or if the cheap tiers get more expensive once the labs are public.

What to watch

  • Whether OpenAI or Anthropic hold the $0.10/$0.50 and $2/$10 tiers once IPO filings put per-call margins on the record.
  • Open-source volume share on OpenRouter in the months after the September 22 launches.
  • Whether DeepSeek or another Chinese provider cuts input prices below GPT-6 Luna's $0.10 per million tokens.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories