Invest1 publisher3 min readPublished
H100 rentals are back to $2.35 an hour, and your AI cost model is stale
GPU rental prices fell more than 60% in 2025. SemiAnalysis data says H100s have since climbed about 40% to $2.35 per GPU-hour, with 12 to 18 month waits for clusters.
The Investor · Invest desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- H100 GPU rental costs climbed roughly 40% between October 2025 and March 2026, with average rates going from $1.70 per GPU-hour to $2.35 per GPU-hour, according to data from SemiAnalysis.
- The report on H100 rental price increases was published by cryptobriefing.com, via intuitionlabs.ai.
- Securing a cluster of H100 chips now involves a wait of 12 to 18 months.
- During much of 2025, GPU rental prices dropped more than 60% from their 2023-2024 peaks.
- The 2025 price decline was driven by aggressive capacity buildouts from Neocloud specialists including CoreWeave and Lambda, alongside expanded offerings from the major hyperscalers.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
The average on-demand price of an Nvidia H100 went from roughly $1.70 per GPU-hour in October 2025 to roughly $2.35 in March 2026, according to SemiAnalysis data reported by Crypto Briefing [1][2]. If you built your inference gross margin on 2025 compute prices, the input has moved about 38% against you and the wait for a new cluster is now 12 to 18 months [3][17].
That is a reversal, not a continuation. Through much of 2025, rental prices fell more than 60% from their 2023-2024 peaks as CoreWeave, Lambda and other specialists poured on capacity and the hyperscalers widened their own offerings [4][5]. Take those two figures at face value and current pricing still sits around 45% below the boom-era peak [18]. This is not a return to 2023 scarcity economics. It is the end of the deflation that a lot of 2025 business plans quietly assumed would continue.
The supply picture explains the turn. Crypto Briefing reports that on-demand H100 capacity is effectively sold out across both neoclouds and hyperscalers, with inference workloads and multi-agent systems pushing requirements past the installed base [6][7]. Clusters as small as 8 nodes, or 64 GPUs, have become hard to procure on short notice [8]. February 2026 alone saw 15% to 20% increases as providers repriced [9]. On-demand rates now span $2.19 to over $4 per GPU-hour depending on provider and whether capacity was reserved ahead of time, an 83% spread for the same silicon [10][19].
Blackwell is not the escape hatch. Lead times for B200 and GB200 deployments have stretched into mid-2026, and most of the available 2026 capacity is already committed [11].
The contract layer is where the real advantage now sits. Some operators are renewing existing H100 agreements at previously agreed legacy rates, and a few have extended commitments through 2028, a four-year term that was nearly unheard of in cloud GPU markets 18 months ago [12][13]. That produces a two-tier cost base: competitors in the same product category paying materially different prices for identical compute. On a modest 64-GPU cluster, the move from $1.70 to $2.35 is about $41.60 an hour, or roughly $364,000 a year at continuous utilisation [20]. That is the difference between two seed-stage runways.
For providers, the same squeeze reads as higher utilisation and better revenue per rack, and anyone holding a favourable long-term Nvidia supply agreement is running a spread trade: cheap wholesale, expensive retail [14][15]. AWS, Google Cloud and Azure can absorb temporary margin compression on balance sheet; smaller neoclouds with uncommitted capacity have more pricing power than their brand would suggest [16].
Two cautions. This is a single-source read on a market with no public order book, and the publisher's own headline claims a 50% surge in six months while its text cites roughly 40% between October and March [21]. Treat the direction as more reliable than the decimal.
What to watch: whether the early mid-2026 stabilisation reported at these higher levels actually holds once Blackwell deliveries land [22]; whether the reserved-to-spot spread widens past $4, which would signal that spot is now a distress channel rather than a discount one [10]; and whether four-year lock-ins spread from the few to the many, because that is the point at which cheap compute stops being available to late entrants at any price.