Product1 publisher3 min readPublished
Clockwork raises $31 million for software that keeps GPU jobs running through hardware failures
Clockwork Systems raised $31 million and launched TorchSnap, a tool that snapshots distributed AI jobs so they can resume after hardware failures. The customer deployments it cites predate the new feature, so its value depends on how long a buyer's own restarts take.
The Product Desk · Product desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Seligman Ventures, Wing Ventures and Premji Invest co-led the round, with NEA and e& Capital returning, and Clockwork has now raised $73 million in total.
- Meta reported a hardware problem every three hours on average during Llama 3's 54-day training run on a 16,384-GPU cluster.
- Clockwork says reloading a job from its last snapshot can take up to 90 minutes, with healthy GPUs idle the whole time and finished work often repeated.
- TorchSnap is Clockwork's third resilience layer, joining the LinkPass network failover tool and the TorchPass GPU migration software.
- LinkedIn runs LinkPass across its entire GPU fleet to route around optical link and switch failures, eliminating thousands of GPU-hours of downtime a month.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- cost On a cluster the size of Meta's, a single restart at Clockwork's 90-minute ceiling leaves about 24,576 GPU-hours idle, and a buyer sets that figure against the software's price.
- decision The no-code install gets jobs restarting where they stopped, but recovering more progress means writing application-level checkpointing, so a rollout team has to price its own engineers' time into the pilot.
- precedent If Patel is right, teams running inference fleets become buyers of restart tooling too, and the market for Clockwork's layers extends past the occasional giant training run.
For the engineer watching a large distributed run, one bad part stops thousands of good ones. The healthy GPUs wait out the reload from the last snapshot, then repeat work they had already finished [4].
Meta's Llama 3 run shows the scale [3]. A hardware problem every three hours for 54 days comes to roughly 432 interruptions [14]. If every one of them hit Clockwork's 90-minute ceiling, recovery would take 648 of the run's 1,296 hours [16]. Clockwork gives 90 minutes as an upper bound [4], so the 648 hours is a worst case. A buyer's case depends on its own median restart time.
Chief executive Suresh Vasudevan sells the fix in throughput terms. "Fault tolerance is a goodput multiplier: it keeps GPUs doing useful work instead of waiting for recovery or repeating work already done," he said [5]. The customer deployments in the announcement are plainer. LinkedIn's is network failover [10]. Together AI offers TorchPass, the GPU migration tool, as a service on its own clusters [11]. WhiteFiber, a neocloud that rents GPUs to enterprises, uses Clockwork's technology to audit and validate cluster reliability before new workloads go into production [12]. None of the three is described as running TorchSnap, the feature launched with the round.
TorchSnap takes multinode snapshots of distributed inference workloads across every node, with no changes to developer code, so a stopped job can restart where it left off [7]. Teams that want to lose less progress can add checkpointing logic at the application level [8]. Underneath is Clockwork's software layer, which sits between the GPUs and the workloads. It keeps the cluster synchronized and uses nanosecond-accurate telemetry to catch failures before they force a full restart [13].
The Meta example is a training run. TorchSnap, by the company's description, snapshots inference workloads [3][7]. SemiAnalysis analyst Dylan Patel argues the problem has spread. "Cluster fault tolerance used to be a training problem, but it is now an inference problem too," he said [9].
Clockwork's larger claim is that many organizations now care less about securing raw GPU capacity and more about getting work out of the clusters they already own [18]. That claim comes from the company, and the support it offers is the customer list above. Investors have put $73 million into the company, $42 million of it before this round [2][17].
The buyer is narrower than the pitch. It is an operator whose jobs span enough GPUs that one fault idles thousands of them [4]. For a team deciding whether to pilot it, I'd draw a two-by-two from the last quarter's incident log. One axis is how often a job dies. The other is the median time to restart. Frequent failures with long restarts is the quadrant where idle GPU-hours pile up and a resilience layer has its clearest case. Rare failures with quick restarts can probably stay on existing checkpoint scripts. In the two mixed quadrants, the failure type picks the tool: link and switch faults are what LinkPass routes around [10], a failed GPU is the TorchPass migration case, and a job that must restart whole is what TorchSnap is for [6].
What to watch
- A named customer publishing measured TorchSnap restart times against its previous checkpoint reloads.
- Clockwork releasing median recovery times alongside the up-to-90-minute figure.
- Together AI or another GPU host adding TorchSnap to the service it already sells on top of TorchPass.