Zürcher Nachrichten - Clockwork.io Introduces A New Class of Fault Tolerance to End Failure-Driven GPU Waste in AI Training

EUR -
AED 4.183233
AFN 72.900796
ALL 91.877285
AMD 414.797448
ANG 2.039348
AOA 1045.665194
ARS 1736.794975
AUD 1.614555
AWG 2.050323
AZN 1.940938
BAM 1.955317
BBD 2.295033
BDT 140.215334
BGN 1.917554
BHD 0.429594
BIF 3432.051477
BMD 1.139068
BND 1.45584
BOB 13.964547
BRL 5.909605
BSD 1.139418
BTN 109.116023
BWP 15.516164
BYN 3.442749
BYR 22325.7403
BZD 2.291733
CAD 1.611725
CDF 2665.420428
CHF 0.943457
CLF 0.027745
CLP 1095.522249
CNY 7.646851
CNH 7.670833
COP 3814.157008
CRC 518.071915
CUC 1.139068
CUP 27.347239
CVE 110.237745
CZK 24.358299
DJF 202.909834
DKK 7.475483
DOP 67.763247
DZD 152.391449
EGP 59.13438
ERN 17.086026
ETB 184.850299
FJD 2.559772
FKP 0.859986
GBP 0.859642
GEL 2.97871
GGP 0.859986
GHS 13.234728
GIP 0.859986
GMD 83.725913
GNF 10021.522331
GTQ 8.701849
GYD 238.411056
HKD 8.934226
HNL 30.583439
HRK 7.533575
HTG 149.123132
HUF 365.202455
IDR 20405.384913
ILS 3.471949
IMP 0.859986
INR 109.142689
IQD 1492.730945
IRR 1565734.922454
ISK 136.984803
JEP 0.859986
JMD 180.27543
JOD 0.807645
JPY 179.158416
KES 147.67349
KGS 99.609484
KHR 4633.85435
KMF 493.217009
KPW 1025.161906
KRW 1543.648431
KWD 0.351563
KYD 0.949565
KZT 504.7852
LAK 25558.681004
LBP 102039.372314
LKR 376.226984
LRD 195.991544
LSL 18.590804
LTL 3.363373
LVL 0.689012
LYD 7.285199
MAD 10.934597
MDL 20.225999
MGA 5030.756222
MKD 61.55978
MMK 2391.810486
MNT 4098.584327
MOP 9.206124
MRU 45.840667
MUR 54.140351
MVR 17.599037
MWK 1975.811511
MXN 20.137822
MYR 4.640683
MZN 72.79829
NAD 18.590804
NGN 1510.837951
NIO 41.929634
NOK 10.830092
NPR 174.585836
NZD 1.992947
OMR 0.439191
PAB 1.139418
PEN 3.868244
PGK 5.076745
PHP 71.01754
PKR 315.746636
PLN 4.37146
PYG 6716.339487
QAR 4.153473
RON 5.27298
RSD 117.420969
RUB 96.076247
RWF 1684.083636
SAR 4.274755
SBD 9.11313
SCR 15.836085
SDG 685.153819
SEK 11.296996
SGD 1.45562
SHP 0.859999
SLE 28.078458
SLL 23885.685201
SOS 651.238991
SRD 42.905863
STD 23576.41575
STN 24.493944
SVC 9.970535
SYP 14810.167401
SZL 18.586405
THB 38.01645
TJS 10.511601
TMT 3.99813
TND 3.373266
TOP 2.742603
TRY 55.748859
TTD 7.750084
TWD 36.141163
TZS 3024.152324
UAH 51.024285
UGX 4462.896617
USD 1.139068
UYU 45.648714
UZS 13485.665874
VES 970.977713
VND 29588.440307
VUV 134.963505
WST 3.127488
XAF 655.957
XAG 0.017716
XAU 0.000265798393
XCD 3.07839
XCG 2.053592
XDR 0.80538
XOF 655.957
XPF 119.331742
YER 269.560946
ZAR 18.57165
ZMK 10252.986409
ZMW 22.227505
ZWL 366.779554
SSP 6507.032816
MXV 2.282055
  • BCC

    1.0400

    77.14

    +1.35%

  • BCE

    -0.3300

    20.97

    -1.57%

  • NGG

    0.2600

    75.49

    +0.34%

  • AZN

    2.0200

    166.58

    +1.21%

  • RIO

    0.0900

    94.56

    +0.1%

  • BTI

    -0.3900

    55.63

    -0.7%

  • BP

    -0.2600

    44.15

    -0.59%

  • GSK

    -0.4100

    49.24

    -0.83%

  • CMSC

    -0.1100

    20.4

    -0.54%

  • JRI

    -0.1500

    11.02

    -1.36%

  • RYCEF

    -0.3600

    19.31

    -1.86%

  • RELX

    0.0100

    33.52

    +0.03%

  • RBGPF

    -0.5900

    65.4

    -0.9%

  • VOD

    0.1300

    16.62

    +0.78%

  • CMSD

    -0.0700

    20.3

    -0.34%

Clockwork.io Introduces A New Class of Fault Tolerance to End Failure-Driven GPU Waste in AI Training
Clockwork.io Introduces A New Class of Fault Tolerance to End Failure-Driven GPU Waste in AI Training

Clockwork.io Introduces A New Class of Fault Tolerance to End Failure-Driven GPU Waste in AI Training

New TorchPass solution addresses a multi-million dollar challenge with AI infrastructure; uses Live GPU Migration to keep large-scale AI training running through hardware failures instead of forcing costly restarts

Text size:

PALO ALTO, CA / ACCESS Newswire / March 11, 2026 / Clockwork.io, the leader in Software-Driven AI Fabrics™- a programmable, vendor-neutral software layer that optimizes large-scale GPU clusters for real-time observability, fault tolerance, and deterministic performance-today announced the general availability of TorchPass Workload Fault Tolerance. This new class of software-driven fault-tolerance eliminates one of the most costly failure modes in large-scale AI training: catastrophic job restarts caused by infrastructure faults.

Delivered as a core capability of the Clockwork.io FleetIQ™ platform, TorchPass applies the principles of Software-Driven AI Fabrics to distributed training, using Live GPU Migration to allow workloads to continue running through GPU failures, network disruptions, driver bugs, and even full node crashes-without checkpoint restarts or lost progress.

"Companies are investing billions in next-gen chips, yet the costs of running distributed AI jobs remains grossly inflated because the ecosystem has accepted failure as a constant," said Suresh Vasudevan, CEO of Clockwork.io. "We built TorchPass to fundamentally reject that premise. Instead of treating failure as inevitable and restarting after the fact, TorchPass makes infrastructure faults invisible to the workload-training continues through failures transparently, in software. For a typical 2,048-GPU deployment, that translates into over $6 million a year in recovered compute. This is what our Software-Driven AI Fabric approach was designed to deliver: fault-tolerant AI infrastructure."

Dylan Patel, Founder and CEO of SemiAnalysis agreed that large-scale training jobs are limited by interruptions.

"As Blackwell clusters roll out with an NVL72 domain, and we look to the future with Rubin Ultra's NVL576 domain, the idea that a single GPU error or network link flap can take down an entire run is totally unacceptable," said Patel. "TorchPass solves a huge challenge with cluster reliability: it provides transparent failover and live workload migration that keeps MFU high, which in turn drives better GPU economics."

Why AI Training Fails at Scale

Distributed AI training remains one of the most failure-prone workloads in modern infrastructure. As cluster sizes grow, fragility increases sharply. Research from Meta FAIR shows that mean time to failure drops to 7.9 hours in a 1,024-GPU cluster and to just 1.8 hours at 16,384 GPUs. This means that for most large, AI-focused enterprises or AI clouds, failure-driven restarts are completely inevitable - making this a major barrier to scaling AI's impact.

Each failure forces training jobs to roll back to the most recent checkpoint, discarding minutes or hours of completed work and wasting additional time on manual intervention, reprovisioning resources and restarting training. These restarts silently cap GPU utilization, making reliability one of the largest hidden costs in AI infrastructure.

TorchPass addresses this problem by proactively addressing costly AI workload failures, solving them before the job stops or needs to restart. Vital for enterprises running large AI workloads and AI clouds alike, TorchPass dramatically improves the reliability of workloads and cluster utilization. For AI clouds, who can now address impacted GPUs while preserving the training run as planned, this translates into better customer SLAs and overall AI cloud economics, improving their ability to protect margin and deliver new models sooner.

"Managing compute output across large-scale GPU clusters is vital to ensuring we're delivering reliable capacity to our customers. By using TorchPass we have the support of a company that focuses on resilience like it is a core business function: it replaces any specific failing GPU and keeps the rest of the job moving, rather than making one small problem impact our large-scale operations," said David Power, CTO of Nscale. "In our evaluation, Live GPU Migration preserved both run continuity and throughput under real fault conditions, which is exactly what you need to deliver predictable time-to-train and a better customer experience at scale."

How Live GPU Migration Works: Reliability Without Restart

TorchPass performs transparent, in-flight migration of impacted training ranks to spare resources when failures occur. TorchPass typically completes recovery in approximately three minutes while the training process continues uninterrupted.

It supports resilience across three failure scenarios:

  • Unplanned migration, handling sudden events such as kernel crashes, power failures, or GPU faults by reconstructing state from healthy replicas

  • Pre-emptive migration, triggered by early warning signals such as rising temperatures or ECC memory errors, enabling controlled migration before a hard failure

  • Planned migration, enabling maintenance, patching, and workload rebalancing without interrupting training

This approach reduces wasted training progress by 95%, cutting lost time from approximately three hours per day to under ten minutes in a 1,024-GPU cluster.

Jordan Nanos, Member of Technical Staff and lead author of ClusterMAX-SemiAnalysis' independent benchmark for large-scale AI training-stress tested Clockwork.io TorchPass and found it delivered leading performance and efficiency for large-scale distributed training, enabling users to reduce checkpointing overhead in training. He shared the following results:

"In our testing, Clockwork.io TorchPass delivered the fastest and most efficient fault-tolerant performance for a gpt-oss-120B training run. We used TorchTitan on a Kubernetes cluster with 64x H200 GPUs. During our testing we measured job completion time (JCT) and Model FLOPs Utilization (MFU) against a standard approach (checkpoint-restart) and the leading open-source fault-tolerant training framework (TorchFT). We simulated multiple hardware failures on the cluster in order to stress test the fault-tolerant training frameworks.

When compared to checkpoint-restart, TorchPass was significantly faster to recover from failures. This reduced overall JCT and maintained high MFU. And when compared to TorchFT, TorchPass had a significantly higher MFU. This reduced overall JCT while also maintaining an equal time to recover from failures.

Using TorchPass also has a downstream effect where it provides users with an opportunity to reduce or even remove checkpointing from their training code. This means larger effective batch sizes, lower risk of out of memory errors (OOMs), and less time spent thinking about storage. For a research organization, this can ultimately mean a faster time to reach their training objective," concluded Nanos.

Measurable Business Impact from Software-Driven Fault-Tolerance

For customers operating large AI clusters, the impact is immediate and measurable. In a typical 2,048-GPU H200 deployment, TorchPass Workload Fault Tolerance delivers over $6 million in annual savings by preventing wasted compute.

These savings come from eliminating hundreds of thousands of GPU-hours that would otherwise be lost to failure-driven restarts, cascading retries, and idle recovery time. By keeping training jobs running through infrastructure faults instead of restarting them, TorchPass converts lost GPU time into productive training, significantly improving the return on GPU investments that today often operate at just 30-50% of theoretical performance.

Enabling the Next Generation of AI Infrastructure

By making reliability a software-defined capability rather than a hardware constraint, TorchPass provides the operational confidence required to deploy next-generation, tightly coupled systems such as NVIDIA GB200 and GB300 NVL72 and future rack-scale systems, where dense architectures amplify the cost of even small failures.

TorchPass builds on Clockwork.io's prior release of Network Fault Tolerance, which applies the same Software-Driven AI Fabric principles to network resilience by transparently rerouting traffic around link failures.

Together, these capabilities form Clockwork.io's Software-Driven AI Fabric, a vendor-neutral software layer spanning network, compute, and storage. As modern AI workloads run on tightly coupled clusters where hundreds or thousands of processors must operate in coordinated lockstep, infrastructure behaves as a single system, where reliability and performance directly determine overall efficiency. By managing this complexity in software, Clockwork.io enables operators to run heterogeneous AI infrastructure as a unified platform-maintaining high utilization, predictable performance, and resilience while preserving the flexibility to evolve hardware and improve the economics of large-scale AI deployments.

To learn more about the launch of TorchPass, visit the Clockwork.io team in-person at NVIDIA GTC from March 16-19, Booth #205, or visit https://clockwork.io.

About Clockwork.io
Clockwork.io pioneers Software-Driven AI Fabrics™, delivering a programmable software layer that makes large-scale AI clusters observable, deterministic, and resilient by design to drive continuous workload progress and peak cluster utilization. Its FleetIQ platform enables enterprises to train, deploy, and serve the world's most demanding AI workloads faster, more reliably, and at lower cost. Companies including Uber, Wells Fargo, DCAI, Nebius, Nscale, and White Fiber trust Clockwork.io to power their AI infrastructure. Learn more at www.clockwork.io.

Media Contact
Dana Trismen
[email protected]
650-269-7478

SOURCE: Clockwork



View the original press release on ACCESS Newswire

T.L.Marti--NZN