Zürcher Nachrichten - Clockwork.io Introduces A New Class of Fault Tolerance to End Failure-Driven GPU Waste in AI Training

EUR -
AED 4.243163
AFN 76.836825
ALL 93.168392
AMD 422.683441
AOA 1059.490664
ARS 1731.895657
AUD 1.635936
AWG 2.081143
AZN 1.964472
BAM 1.956127
BBD 2.327789
BDT 142.803865
BHD 0.435867
BIF 3452.176908
BMD 1.155388
BND 1.478954
BOB 13.753358
BRL 5.87804
BSD 1.155693
BTN 110.085873
BWP 15.570074
BYN 3.435989
BYR 22645.606405
BZD 2.324398
CAD 1.610755
CDF 2614.066317
CHF 0.933808
CLF 0.026824
CLP 1055.723909
CNY 7.7961
CNH 7.793635
COP 3642.60356
CRC 524.689954
CUC 1.155388
CUP 30.617784
CVE 110.283907
CZK 24.250092
DJF 205.808847
DKK 7.47569
DOP 67.38126
DZD 153.592261
EGP 57.747336
ERN 17.330821
ETB 186.880539
FJD 2.553243
FKP 0.857011
GBP 0.85579
GEL 3.016025
GGP 0.857011
GHS 13.568326
GIP 0.857011
GMD 84.917311
GNF 10149.74009
GTQ 8.818397
GYD 242.045687
HKD 9.064453
HNL 30.979177
HRK 7.536369
HTG 151.115254
HUF 363.604111
IDR 20498.202116
ILS 3.466245
IMP 0.857011
INR 110.109235
IQD 1514.05302
IRR 1589294.076063
ISK 142.216689
JEP 0.857011
JMD 183.540672
JOD 0.819128
JPY 183.187355
KES 149.507516
KGS 101.038588
KHR 4685.803343
KMF 492.195074
KRW 1637.520047
KWD 0.3569
KYD 0.963165
KZT 538.813048
LAK 26101.474886
LBP 103496.143572
LKR 387.344731
LRD 208.609375
LSL 18.700163
LTL 3.411561
LVL 0.698883
LYD 7.365231
MAD 10.777548
MDL 20.086965
MGA 4949.848611
MKD 61.530283
MMK 2425.794575
MNT 4152.757665
MOP 9.339844
MRU 46.355747
MUR 54.303359
MVR 17.850882
MWK 2004.078274
MXN 19.818137
MYR 4.726342
MZN 73.835072
NAD 18.700163
NGN 1573.280637
NIO 42.527107
NOK 10.987336
NPR 176.136435
NZD 1.961554
OMR 0.444245
PAB 1.155698
PEN 3.903852
PGK 5.109876
PHP 70.211802
PKR 320.85972
PLN 4.299164
PYG 6879.149604
QAR 4.213104
RON 5.244535
RSD 117.352589
RUB 95.451742
RWF 1697.791072
SAR 4.326647
SBD 9.319009
SCR 17.004638
SDG 693.805905
SEK 10.965222
SGD 1.478256
SLE 28.420632
SOS 660.501804
SRD 43.750504
STD 23914.200576
STN 24.504095
SVC 10.112734
SZL 18.697125
THB 38.145135
TJS 10.673415
TMT 4.055412
TND 3.387866
TRY 55.125778
TTD 7.839344
TWD 37.238734
TZS 3058.893435
UAH 51.843388
UGX 4304.71938
USD 1.155388
UYU 46.557982
UZS 13790.304557
VES 873.198868
VND 30219.175282
VUV 137.91453
WST 3.158504
XAF 656.069478
XAG 0.018024
XAU 0.000266
XCD 3.122494
XCG 2.082948
XDR 0.815502
XOF 656.069478
XPF 119.331742
YER 275.449418
ZAR 18.687033
ZMK 10399.872862
ZMW 21.617706
ZWL 372.034491
  • NGG

    -1.1200

    79.76

    -1.4%

  • BCC

    -1.7000

    84.9

    -2%

  • GSK

    -1.0800

    51.88

    -2.08%

  • RIO

    0.3650

    101.465

    +0.36%

  • BTI

    -1.6600

    57.67

    -2.88%

  • JRI

    -0.0400

    12.77

    -0.31%

  • BP

    0.9400

    42.57

    +2.21%

  • BCE

    -0.2550

    22.495

    -1.13%

  • RYCEF

    -0.0300

    20.97

    -0.14%

  • RBGPF

    0.8600

    70.6

    +1.22%

  • CMSD

    -0.0500

    21.77

    -0.23%

  • CMSC

    -0.2037

    21.5401

    -0.95%

  • RELX

    0.1050

    35.625

    +0.29%

  • AZN

    0.2200

    161.64

    +0.14%

  • VOD

    -0.3750

    15.815

    -2.37%

Clockwork.io Introduces A New Class of Fault Tolerance to End Failure-Driven GPU Waste in AI Training
Clockwork.io Introduces A New Class of Fault Tolerance to End Failure-Driven GPU Waste in AI Training

Clockwork.io Introduces A New Class of Fault Tolerance to End Failure-Driven GPU Waste in AI Training

New TorchPass solution addresses a multi-million dollar challenge with AI infrastructure; uses Live GPU Migration to keep large-scale AI training running through hardware failures instead of forcing costly restarts

Text size:

PALO ALTO, CA / ACCESS Newswire / March 11, 2026 / Clockwork.io, the leader in Software-Driven AI Fabrics- a programmable, vendor-neutral software layer that optimizes large-scale GPU clusters for real-time observability, fault tolerance, and deterministic performance-today announced the general availability of TorchPass Workload Fault Tolerance. This new class of software-driven fault-tolerance eliminates one of the most costly failure modes in large-scale AI training: catastrophic job restarts caused by infrastructure faults.

Delivered as a core capability of the Clockwork.io FleetIQ platform, TorchPass applies the principles of Software-Driven AI Fabrics to distributed training, using Live GPU Migration to allow workloads to continue running through GPU failures, network disruptions, driver bugs, and even full node crashes-without checkpoint restarts or lost progress.

"Companies are investing billions in next-gen chips, yet the costs of running distributed AI jobs remains grossly inflated because the ecosystem has accepted failure as a constant," said Suresh Vasudevan, CEO of Clockwork.io. "We built TorchPass to fundamentally reject that premise. Instead of treating failure as inevitable and restarting after the fact, TorchPass makes infrastructure faults invisible to the workload-training continues through failures transparently, in software. For a typical 2,048-GPU deployment, that translates into over $6 million a year in recovered compute. This is what our Software-Driven AI Fabric approach was designed to deliver: fault-tolerant AI infrastructure."

Dylan Patel, Founder and CEO of SemiAnalysis agreed that large-scale training jobs are limited by interruptions.

"As Blackwell clusters roll out with an NVL72 domain, and we look to the future with Rubin Ultra's NVL576 domain, the idea that a single GPU error or network link flap can take down an entire run is totally unacceptable," said Patel. "TorchPass solves a huge challenge with cluster reliability: it provides transparent failover and live workload migration that keeps MFU high, which in turn drives better GPU economics."

Why AI Training Fails at Scale

Distributed AI training remains one of the most failure-prone workloads in modern infrastructure. As cluster sizes grow, fragility increases sharply. Research from Meta FAIR shows that mean time to failure drops to 7.9 hours in a 1,024-GPU cluster and to just 1.8 hours at 16,384 GPUs. This means that for most large, AI-focused enterprises or AI clouds, failure-driven restarts are completely inevitable - making this a major barrier to scaling AI's impact.

Each failure forces training jobs to roll back to the most recent checkpoint, discarding minutes or hours of completed work and wasting additional time on manual intervention, reprovisioning resources and restarting training. These restarts silently cap GPU utilization, making reliability one of the largest hidden costs in AI infrastructure.

TorchPass addresses this problem by proactively addressing costly AI workload failures, solving them before the job stops or needs to restart. Vital for enterprises running large AI workloads and AI clouds alike, TorchPass dramatically improves the reliability of workloads and cluster utilization. For AI clouds, who can now address impacted GPUs while preserving the training run as planned, this translates into better customer SLAs and overall AI cloud economics, improving their ability to protect margin and deliver new models sooner.

"Managing compute output across large-scale GPU clusters is vital to ensuring we're delivering reliable capacity to our customers. By using TorchPass we have the support of a company that focuses on resilience like it is a core business function: it replaces any specific failing GPU and keeps the rest of the job moving, rather than making one small problem impact our large-scale operations," said David Power, CTO of Nscale. "In our evaluation, Live GPU Migration preserved both run continuity and throughput under real fault conditions, which is exactly what you need to deliver predictable time-to-train and a better customer experience at scale."

How Live GPU Migration Works: Reliability Without Restart

TorchPass performs transparent, in-flight migration of impacted training ranks to spare resources when failures occur. TorchPass typically completes recovery in approximately three minutes while the training process continues uninterrupted.

It supports resilience across three failure scenarios:

  • Unplanned migration, handling sudden events such as kernel crashes, power failures, or GPU faults by reconstructing state from healthy replicas

  • Pre-emptive migration, triggered by early warning signals such as rising temperatures or ECC memory errors, enabling controlled migration before a hard failure

  • Planned migration, enabling maintenance, patching, and workload rebalancing without interrupting training

This approach reduces wasted training progress by 95%, cutting lost time from approximately three hours per day to under ten minutes in a 1,024-GPU cluster.

Jordan Nanos, Member of Technical Staff and lead author of ClusterMAX-SemiAnalysis' independent benchmark for large-scale AI training-stress tested Clockwork.io TorchPass and found it delivered leading performance and efficiency for large-scale distributed training, enabling users to reduce checkpointing overhead in training. He shared the following results:

"In our testing, Clockwork.io TorchPass delivered the fastest and most efficient fault-tolerant performance for a gpt-oss-120B training run. We used TorchTitan on a Kubernetes cluster with 64x H200 GPUs. During our testing we measured job completion time (JCT) and Model FLOPs Utilization (MFU) against a standard approach (checkpoint-restart) and the leading open-source fault-tolerant training framework (TorchFT). We simulated multiple hardware failures on the cluster in order to stress test the fault-tolerant training frameworks.

When compared to checkpoint-restart, TorchPass was significantly faster to recover from failures. This reduced overall JCT and maintained high MFU. And when compared to TorchFT, TorchPass had a significantly higher MFU. This reduced overall JCT while also maintaining an equal time to recover from failures.

Using TorchPass also has a downstream effect where it provides users with an opportunity to reduce or even remove checkpointing from their training code. This means larger effective batch sizes, lower risk of out of memory errors (OOMs), and less time spent thinking about storage. For a research organization, this can ultimately mean a faster time to reach their training objective," concluded Nanos.

Measurable Business Impact from Software-Driven Fault-Tolerance

For customers operating large AI clusters, the impact is immediate and measurable. In a typical 2,048-GPU H200 deployment, TorchPass Workload Fault Tolerance delivers over $6 million in annual savings by preventing wasted compute.

These savings come from eliminating hundreds of thousands of GPU-hours that would otherwise be lost to failure-driven restarts, cascading retries, and idle recovery time. By keeping training jobs running through infrastructure faults instead of restarting them, TorchPass converts lost GPU time into productive training, significantly improving the return on GPU investments that today often operate at just 30-50% of theoretical performance.

Enabling the Next Generation of AI Infrastructure

By making reliability a software-defined capability rather than a hardware constraint, TorchPass provides the operational confidence required to deploy next-generation, tightly coupled systems such as NVIDIA GB200 and GB300 NVL72 and future rack-scale systems, where dense architectures amplify the cost of even small failures.

TorchPass builds on Clockwork.io's prior release of Network Fault Tolerance, which applies the same Software-Driven AI Fabric principles to network resilience by transparently rerouting traffic around link failures.

Together, these capabilities form Clockwork.io's Software-Driven AI Fabric, a vendor-neutral software layer spanning network, compute, and storage. As modern AI workloads run on tightly coupled clusters where hundreds or thousands of processors must operate in coordinated lockstep, infrastructure behaves as a single system, where reliability and performance directly determine overall efficiency. By managing this complexity in software, Clockwork.io enables operators to run heterogeneous AI infrastructure as a unified platform-maintaining high utilization, predictable performance, and resilience while preserving the flexibility to evolve hardware and improve the economics of large-scale AI deployments.

To learn more about the launch of TorchPass, visit the Clockwork.io team in-person at NVIDIA GTC from March 16-19, Booth #205, or visit https://clockwork.io.

About Clockwork.io
Clockwork.io pioneers Software-Driven AI Fabrics™, delivering a programmable software layer that makes large-scale AI clusters observable, deterministic, and resilient by design to drive continuous workload progress and peak cluster utilization. Its FleetIQ platform enables enterprises to train, deploy, and serve the world's most demanding AI workloads faster, more reliably, and at lower cost. Companies including Uber, Wells Fargo, DCAI, Nebius, Nscale, and White Fiber trust Clockwork.io to power their AI infrastructure. Learn more at www.clockwork.io.

Media Contact
Dana Trismen
[email protected]
650-269-7478

SOURCE: Clockwork



View the original press release on ACCESS Newswire

T.L.Marti--NZN