NVIDIA GPUs' Warp Divergence Cost Model Hasn't Changed Since Pascal, But the ISA Has
Characterizing Warp Divergence from Pascal to Blackwell

A new study tests the assumption that NVIDIA GPUs handle warp divergence uniformly since Volta's Independent Thread Scheduling. Using microbenchmarks, hardware counters, and SASS analysis across Pascal, Ampere, Hopper, and Blackwell, the authors find that divergent paths serialize linearly with path count, with no super-linear reconvergence penalty. Warp execution efficiency drops as 32/k, independent of occupancy, and predication removes the cost. However, the compiler-emitted reconvergence machinery has evolved: Pascal used a per-warp SSY/SYNC stack, while later generations use barrier registers. Deferred reconvergence cases fell from 29 on Ampere to 2 on Blackwell, which also introduces new convergence-barrier classifications and explicit partial-mask warp synchronization. The new barrier class appears to be a static compiler classification with no observable runtime effect.
Thus, divergence retains a stable and predictable performance cost even as NVIDIA's control-flow ISA and reconvergence mechanisms continue to evolve.