From Zero to 94% of FlashAttention-4: A B200 Attention Kernel Built in 60 Diagrams
B200 Attention Kernel from Scratch to Near-SOTA in 60 Diagrams

This hands-on guide builds a dense B200 attention kernel from scratch in CUDA and PTX, progressing from a naive baseline to 94.4% of FlashAttention-4 performance on 4K, 8K, and 16K sequence lengths. The author uses 60 diagrams to explain each optimization step, covering Blackwell-specific features like tcgen05 MMA and TMEM, and addresses the softmax bottleneck caused by asymmetric hardware scaling. The final kernel is plugged into a video-generation model as a capstone project. All code is available on GitHub.
It's one of the hardest kernels out there, running on the latest hardware, so it'll be fun.