Why the Elegant Muon Optimizer Fails on Transformer QK Stability

From Muon to Gradient Clipping: Some Thoughts on QK Stability

10Eridanus2💬 0

I explore why the theoretically elegant Muon optimizer causes training collapse when applied to Transformer Query and Key matrices. While Muon constrains updates in function space, its full-rank nature conflicts with the bilinear coupling of attention, leading to spectral norm explosions. I trace this failure from first principles and propose a new approach that directly constrains the attention score product rather than individual weights.

"The elegance of the theory runs into a wall in practice."

More from this day · 2026-07-19