Bringing PyTorch Monarch to AMD GPUs for Fault-Tolerant Training

Bringing PyTorch Monarch to AMD GPUs for Fault-Tolerant Training

We have extended PyTorch Monarch to AMD Instinct GPUs with ROCm, enabling single-controller distributed training that dynamically recovers from hardware failures. By integrating with TorchTitan and TorchFT, our system allows healthy nodes to continue training while failed nodes restart, eliminating the need for costly full-cluster checkpoints. This breakthrough ensures stable, large-scale AI infrastructure where reliability matches performance across diverse environments like SLURM and Kubernetes.

At this scale, hardware failures are not exceptional events—they are expected.
  1. mountainriver

    Love monarch, amazing primitives and feels so much lighter than Ray.

    Another awesome add on their part

  2. bicepjai

    So are we at a point where we can’t train llms for fun at home on less expensive AMD cards ?

  3. fefe23

    benchmqarks?

More from this day

2026-07-25