Why the Elegant Muon Optimizer Fails on Transformer QK Stability
From Muon to Gradient Clipping: Some Thoughts on QK Stability
I explore why the theoretically elegant Muon optimizer causes training collapse when applied to Transformer Query and Key matrices. While Muon constrains updates in function space, its full-rank nature conflicts with the bilinear coupling of attention, leading to spectral norm explosions. I trace this failure from first principles and propose a new approach that directly constrains the attention score product rather than individual weights.
"The elegance of the theory runs into a wall in practice."