Skip to content

Build1 publisher2 min readPublished

Clockwork raises $31 million to keep failed GPU jobs running without a restart

Clockwork.io raised $31 million as LinkedIn runs its LinkPass failover across its AI infrastructure and Together AI sells TorchPass job recovery. Restart-free recovery now runs on paying GPU clusters, though the savings figures all come from customers.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Clockwork raises $31 million to keep failed GPU jobs running without a restart
Generated illustration

What happened

  • LinkPass reroutes traffic around a failed network link, while TorchPass moves a job off a failed GPU onto a healthy one and keeps its training progress.
  • LinkedIn infrastructure CTO Raghu Hiremagalur says LinkPass prevents tens of thousands of GPU-hours of downtime a month across LinkedIn's AI infrastructure.
  • The October 5 release adds TorchPass multi-node snapshots meant to capture a distributed job's state without any change to training code.
  • WhiteFiber, already a customer, says it is expanding Clockwork's software across its global GPU-as-a-service network.
  • Premji Invest, Wing Venture Capital and Seligman Ventures led the round, which Clockwork says brings its total raised to $73 million.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost The cluster owner pays for every restart: at the release's 90-minute worst case, each healthy GPU in a failed job waits, and then the job repeats its lost work.
  • decision Operators sizing a purchase have to bring their own failure rates and restart times, because the savings evidence on offer is customer-reported and tied to each customer's cluster.
  • capability Teams renting GPUs from Together AI can buy job recovery from their provider instead of building it into their own training stack.

Clockwork prices its products against checkpoint-restart. According to the release, recovery on that path can take as long as 90 minutes. The healthy GPUs wait, and the job then repeats everything since its last saved state [4]. For the failure rate, the release points to what Meta reported from its Llama 3 training: over a 54-day run on 16,384 GPUs, an unexpected interruption hit about once every three hours [3].

Those two numbers set the scale. Fifty-four days at one interruption every three hours comes to about 432 interruptions [17]. One 90-minute stall across 16,384 GPUs idles 24,576 GPU-hours [18]. Both figures use the release's worst case and describe the size of the problem.

Hiremagalur's monthly figure is the one savings number in the announcement [2]. RuntimeWire, which reported the round, describes the savings figures as customer-reported [12]. Set against a 24,576-GPU-hour stall, LinkedIn's claim equals roughly one to four worst-case restarts a month on a job the size of Llama 3 [19]. "Tens of thousands" is a wide band for a resource billed by the hour. The figure also covers LinkPass alone, so it measures network-link failures. Failed GPUs are TorchPass's job [7].

The number transfers to another cluster only under three conditions. The link-failure rate has to be similar. Each failure has to stall a similar number of GPUs. And the restart path it replaced has to be as slow as the buyer's own.

The claim I would test first is the snapshot one, because it sets the adoption cost [8]. If snapshots really need no training-code changes, adoption is an infrastructure install and the training loop stays as it is. If they do need changes, the product competes with checkpointing code that teams have already written. The asynchronous checkpoint feature is aimed at reinforcement learning. It pushes updated weights to inference replicas, so they wait less and generate training examples from fresher weights [9].

Clockwork came to recovery from clock synchronization. It started in 2018 as TickTock Networks, and its technology is built on the Huygens clock-synchronization system from co-founder Yilong Geng's Stanford research [14]. Its $21 million Series A in March 2022, led by New Enterprise Associates, came while it was selling network visibility and clock synchronization [15]. The new money goes to fault-tolerance deployments across training, inference and reinforcement learning, including through cloud partners [16]. The release does not provide contract values or a breakdown of deployments by product [13].

What to watch

  • A customer publishing its failure rate and old restart time alongside its GPU-hour savings, so the figure can be compared across clusters.
  • Pricing or recovery-time data from Together AI's TorchPass service, the first place outside users can test the no-code-change snapshot claim.
  • Named cloud partners for the fault-tolerance expansion Clockwork says the round will fund.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories