Bringing PyTorch Monarch to AMD GPUs for Fault-Tolerant Training

We have extended PyTorch Monarch to AMD Instinct GPUs with ROCm, enabling single-controller distributed training that dynamically recovers from hardware failures. By integrating with TorchTitan and TorchFT, our system allows healthy nodes to continue training while failed nodes restart, eliminating the need for costly full-cluster checkpoints. This breakthrough ensures stable, large-scale AI infrastructure where reliability matches performance across diverse environments like SLURM and Kubernetes.
At this scale, hardware failures are not exceptional events—they are expected.