Invest1 distinct publisher3 min readUpdated
Silicon Data figures reported by the FT put the one-month drop at close to 25%. The pressure is coming from DeepSeek and Moonshot, not from OpenAI and Anthropic fighting each other.
The Investor · Invest desk

Compiled by The InvestorSomething wrong?How this is made
Average prices for AI inference from leading US labs fell nearly 25% between mid-July and mid-August, according to Silicon Data analysis reported by the Financial Times [1]. If your cost model, your gross-margin deck, or your build-versus-buy memo was priced off mid-2026 rate cards, it is now overstating your inference bill by roughly a third [1].
The mechanics are specific rather than atmospheric. OpenAI moved on July 30, cutting two of the three GPT-5.6 tiers [2]: the mid-range Luna by 80% [3], and Terra by 20% [4]. Flagship Sol held its price, with OpenAI accelerating its performance options at the same price point [5]. Anthropic launched Claude Opus 5 at half the price of its predecessor, Fable 5 [6] - a deeper cut than Terra's but well short of Luna's [3].
The asymmetry is the tell. Luna now sells for one-fifth of its old price, while Terra keeps four-fifths of its own [2]. Cryptobriefing reads OpenAI's decision to hold Sol as a tiered strategy in which the best model defends margin and the cheaper tiers compete for volume [11]. That is the sane interpretation, and it also tells buyers where the negotiating room is: not at the frontier.
Crucially, this is not two US labs bidding each other down. The source attributes the pressure primarily to Chinese providers, with DeepSeek and Moonshot leading, offering models that match or closely approach US benchmarks at dramatically lower prices [7][8]. That matters for how long the discounting lasts. A domestic price war can end with a truce; a cost-structure competitor with different capital expectations does not negotiate.
Set the month against the decade. GPT-4-class capability cost more than $20 per million tokens in late 2022 and less than $1 by mid-2026, a decline of more than 95% in about three and a half years [9] - a trend rate of at least 57% a year [4]. One month at 25% off, if repeated, would compound to about a 97% annual decline [5]. Nobody should model that as the run rate, but it does say the past month was several times steeper than trend, and steepness is what breaks planning assumptions.
For buyers, this is straightforwardly good: cheaper inference removes a barrier to adoption and lets workloads that did not pencil out clear the bar [13]. For the sell side it is harder. US labs have spent billions on inference infrastructure justified by revenue projections that assumed particular price levels, and when prices fall faster than usage grows the return on that spending stretches [12]. Valuations built partly on the premise that inference stays a high-margin business need revising when revenue per query drops a quarter in a month [10].
What to watch: whether the next monthly read from Silicon Data shows the cut holding or partly clawed back; whether Sol's price survives the next DeepSeek or Moonshot release [5][7]; and whether token volumes at the labs grow faster than the price declines, which is the only way the infrastructure math closes [12]. On the buy side, check whether your vendor contracts pass cuts through automatically or require you to ask.
Ranked by verification strength, evidence, and original report placement.
Average prices for AI inference from leading US labs dropped nearly 25% between mid-July and mid-August, according to Silicon Data analysis reported by the Financial Times.
OpenAI cut prices on two of its three GPT-5.6 model tiers on July 30.
GPT-5.6 Luna, OpenAI's mid-range offering, received an 80% price reduction.
GPT-5.6 Terra received a 20% price reduction.
Flagship GPT-5.6 Sol held its price steady, while OpenAI accelerated its performance options at the same price point.
GPT-4-class performance cost over $20 per million tokens in late 2022 and less than $1 per million tokens by mid-2026 across various providers, a decline of more than 95% in roughly three and a half years.
Distinct publishers with included, body-backed reporting in this cluster.
cryptobriefing.com
1 article · August 16, 2026
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single third-hand source; price mechanics stated, magnitudes unverifiable
Everything rests on one crypto trade article that relays a Financial Times report of a Silicon Data analysis. The per-tier cuts and the ~25% aggregate decline are stated concretely, but no vendor pricing page, index methodology, absolute per-token price, or benchmark result is supplied, and the model names cannot be checked against any primary artifact in the cluster. Interpretive and forward-looking claims about valuations and infrastructure returns carry no financial data at all.
Supplier-side price moves observed; no usage evidence
Three concrete pricing observations are supplied: OpenAI's July 30 tier cuts, Anthropic's Opus 5 launch at half of Fable 5, and the aggregate index decline. All are vendor pricing actions rather than demand signals. No customer deployment, token-volume, or usage disclosure appears anywhere in the source, so the claim that cheaper inference unlocks enterprise workloads has no measured uptake behind it.
One month of third-hand data carrying valuation-scale conclusions
The reported price cuts are plausible and specifically described, but the framing stretches well past them: a single month of an aggregate index is treated as a durable revenue-per-query shock, list-price cuts are equated with margin compression, and the conclusion reaches to repricing the whole AI supply chain. The article's own long-run figure (>95% over ~3.5 years, roughly 57% a year) is far slower than the annualized reading of a 25% monthly drop, which shows the monthly number is being read as more dramatic than the trend supports.
Vendor-interested pricing moves relayed by ad-supported trade press
The underlying actions are competitive commercial moves by the vendors themselves, and the article's own reading is that OpenAI protected flagship margin while discounting lower tiers to defend volume against Chinese entrants, so the price signals are strategically motivated rather than neutral. The relaying publisher is a trade/crypto outlet with attention incentives, and the index provider Silicon Data sells market data. No disclosures, sponsorships, or conflicts are stated in the source, so this is scored on the claim structure rather than on declared interests.
Low: directionally credible, numerically unverified
Confidence is capped by a single-publisher cluster, a two-step attribution chain, and unverifiable product names. The direction of travel — falling inference prices with cheaper tiers cut hardest — is internally consistent and matches the long-run trajectory the article reports, which supports directional trust. The specific magnitudes and every financial-consequence claim would need vendor price pages, the Silicon Data methodology, and volume data before being relied on.
Follow any of these and your For You feed starts watching them — no settings page required.
product
A 27B laptop model scores like a rented one, and thinks three times as hard to do it1 distinct publisher
build
Model provenance now arrives through the billing layer, not the vendor contract2 distinct publishers
leadership
The AI bill nobody reconciles: cost per finished task, not per million tokens1 distinct publisher
invest
Korea cuts one of four sovereign AI teams, and usability did the cutting1 distinct publisher