Clockwork raises $31M as LinkedIn deploys LinkPass across its AI infrastructure
Clockwork.io CEO Suresh Vasudevan says a conversation with co-founder Balaji Prabhakar drew him back to startups; Together AI is commercializing TorchPass on its GPU clusters.
By Ryan Merket · Published
Primary source: PR Newswire
Why it matters
Clockwork is selling recovery software as a way to get more useful work from costly GPU clusters. LinkedIn's LinkPass deployment and Together AI's TorchPass service are named production uses, while the savings figures remain customer-reported.

Clockwork.io, the AI infrastructure startup co-founded by Stanford professor Balaji Prabhakar, has raised $31 million as LinkedIn deploys its LinkPass network-failure technology across its AI infrastructure and Together AI sells TorchPass workload recovery on GPU clusters. The October 5th announcement also introduces new TorchPass features designed to preserve distributed AI jobs through failures without changes to training code.
For CEO Suresh Vasudevan, the round marks a return to startup building after a career spent scaling infrastructure companies. In a post about joining Clockwork, Vasudevan wrote that he thought he was done with startups after six years at Sysdig, 17 years at startups and a decade at NetApp. He had planned to focus on board work, mentoring and perhaps teaching. One conversation with Clockwork co-founder Balaji Prabhakar changed that plan.
Vasudevan previously led Nimble Storage from startup through its IPO and acquisition by HPE, and was CEO of Omneon before its acquisition by Harmonic. At Clockwork, he is taking that operating experience into a problem that lands squarely in infrastructure economics: when one failed GPU or network link stalls a large AI job, operators can pay for idle hardware and work that must be repeated.
From clock synchronization to AI recovery
Clockwork says it was founded in 2018 as TickTock Networks, renamed Clockwork Systems in 2021, and debuted publicly as Clockwork.io in 2022. Co-founder Yilong Geng's Stanford research produced the Huygens clock-synchronization system that underpins the company's technology, according to Clockwork's leadership profile. Co-founder Prabhakar is a Stanford computer science professor; he and co-founder Deepak Merugu previously built Urban Engines, acquired by Google in 2016.
Clockwork announced a $21 million Series A in March 2022, led by New Enterprise Associates, when it was selling network visibility and clock-synchronization technology. The new round, led jointly by Premji Invest, Wing Venture Capital and Seligman Ventures, includes existing investors NEA and e& Capital. Clockwork says the financing takes its total raised to $73 million. It plans to use the money to expand fault-tolerance deployments across training, inference and reinforcement learning, including through cloud partners.
That shift in product focus tracks a change in the infrastructure problem Clockwork is selling against. Its LinkPass software reroutes traffic when a network link fails. TorchPass moves a job from a failed GPU to a healthy one, preserving training progress. The new multi-node snapshots are intended to capture the state of a distributed job across its nodes without requiring changes to training code. A separate asynchronous checkpoint feature sends updated model weights to inference replicas during reinforcement learning, so those replicas can generate training examples without waiting as long or using stale weights.

The commercial proof points are consequential, but the headline metric remains a customer claim. LinkedIn infrastructure CTO Raghu Hiremagalur says in the announcement that LinkPass prevents tens of thousands of GPU-hours of downtime each month across LinkedIn's AI infrastructure. Together AI says it is commercializing TorchPass as a service on its GPU clusters. WhiteFiber, an existing customer, says it is expanding Clockwork's software across its global GPU-as-a-service network. The release does not provide contract values or a breakdown of deployments by product.
Selling productive GPU time
Clockwork's pitch is that keeping a job running is better than waiting for a restart. Its release cites Meta's report of roughly one unexpected interruption every three hours during a 54-day Llama 3 training run on 16,384 GPUs. It says conventional checkpoint recovery can take as long as 90 minutes, leaving healthy GPUs waiting and forcing the job to repeat work since its last saved state. Those figures describe the scale of the problem; they do not establish how much time Clockwork saves across its customers.
A software layer that works across hardware and cloud providers could help enterprise AI teams and GPU cloud operators get more useful work from existing clusters. Together AI's decision to offer TorchPass as a service also gives Clockwork a route to customers through a cloud provider, alongside direct enterprise deployments.
Vasudevan's own explanation of why he joined puts the bet in terms of communication as the new Moore's law: as AI jobs spread across more accelerators, coordination between machines can become the constraint even when the GPUs themselves are available. Clockwork's founders built around precise timing and distributed systems; Vasudevan is now leading the effort to turn that technical foundation into infrastructure operators can deploy and pay for.
Clockwork CEO Suresh Vasudevan put the operating case plainly in the announcement: "Failures are inevitable at AI scale. We should not lose productive hours of work because of them." The $31 million round will support Clockwork's plan to expand fault-tolerance deployments across training, inference and reinforcement learning, including through cloud partners.