Skip to content

Product1 publisher3 min readPublished

Thunder Compute's $13M rests on a number: enterprise GPUs run at 5% to 20% utilization

If GPU virtualization works as claimed, part of the scarcity premium buyers are paying is not a supply shortage but idle time they already own.

The Product Desk · Product desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying Thunder Compute's $13M rests on a number: enterprise GPUs run at 5% to 20% utilization
Photo: thundercompute.com

What happened

  • Thunder Compute announced it has raised $13 million in early funding to help GPU cloud providers squeeze out compute capacity that sits idle and wasted at the long end of workload cycles.
  • Because GPUs are traditionally allocated as bare-metal resources dedicated to individual workloads, the chips can spend much of their reserved time sitting idle waiting for work.
  • Co-founder and CEO Carl Peterson said the objective is to build the "VMware for GPUs" - virtualizing them so the underlying hardware disappears, much like storage or central processing units.
  • According to the CastAI 2026 State of Kubernetes Optimization Report, enterprise GPUs sit idle, averaging around 5% to 20% utilization.
  • Peterson said much of the underutilization comes from how GPUs are reserved: they are allocated continuously to workloads regardless of whether they are actually being used.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

Thunder Compute said it has raised $13 million in early funding to help GPU cloud providers recover the compute capacity that sits idle at the long end of workload cycles [1]. The interesting figure is not the raise but the one underneath it: the CastAI 2026 State of Kubernetes Optimization Report, as cited by SiliconANGLE, puts average enterprise GPU utilization at roughly 5% to 20% [4].

Read that range the other way and 80% to 95% of reserved GPU time produces nothing [17]. Perfect scheduling of a fleet running at those rates would be worth between 5 and 20 times its current output [18]. Nobody hits that ceiling, but the gap between it and the current price of an H-class hour is where the argument lives. If the constraint is partly allocation rather than fabrication, then some of the premium buyers pay for scarce accelerators is a scheduling failure they are financing themselves.

The mechanism is unglamorous. GPUs are traditionally handed out as bare-metal resources dedicated to a single workload, so expensive chips spend much of their reserved time waiting for work [2]. Co-founder and Chief Executive Carl Peterson told SiliconANGLE that the underutilization follows from the reservation model: capacity is allocated continuously whether or not it is being used [6]. Thunder's software separates a workload's access to a GPU from the specific hardware serving it, treating GPUs as network resources reachable across the data center, with the company sitting between the developer and the cloud provider [7][8]. Peterson's framing is "VMware for GPUs" - virtualize the chip until the hardware disappears, the way storage and CPUs already have [3]. He added that the virtualization is meant to be invisible, and that the goal is for developers not to care [9].

That matters for who buys it. Peterson said the economic benefit goes primarily to the cloud provider or enterprise that purchased the GPUs, with the possibility of passing some savings through as lower prices [10]. This is a balance-sheet product sold to asset owners, not a developer tool.

The evidence is thinner than the thesis. Thunder has supplied compute to more than 10,000 users on its own cloud of virtualized GPUs [11], which Peterson described as effectively selling to itself while acting as the cloud provider [13]. Two enterprises are piloting the software and he could not name any customers [12]. He said some customers have seen gains of four times or more, while cautioning that the company cannot promise that to everyone because results depend on the workload and its existing utilization [14]. Four times sits comfortably below the 5x-to-20x arithmetic ceiling implied by the utilization range, which makes it plausible rather than remarkable [19]. The software has been in development for four years, and the Series A is framed as the shift from proving it on Thunder's own cloud to running it on fleets other people own [15].

Watch three things. Whether the two pilots convert into named references on third-party fleets, since a virtualization layer that works on your own homogeneous cloud is a weaker claim than one that survives someone else's [12][15]. Whether any provider actually lowers prices rather than banking the recovered capacity as margin [10]. And the hiring: Peterson said the funding will go partly to systems researchers and to engineers supporting enterprise GPU deployments, which is the cost structure of a company that expects each large fleet to be its own integration problem [16].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories