Squeezing 85.3 GFLOPS from a Single AMD Zen 3 Core with C++ Intrinsics
85.3 GFlops: Optimizing FP32 Matrix Multiplication on a Single AMD Zen 3 Core
I systematically tested 28 optimization configurations for FP32 matrix multiplication on a single AMD Zen 3 core using C++ intrinsics. By fine-tuning cache blocking, register usage, and FMA chaining, my best model achieved 85.30 GFLOPS, reaching 63.5% of the theoretical peak and matching performance from libraries like AMD AOCL and OpenBLAS.
"Using non-temporal stores was catastrophic, dropping performance to just 1.24 GFLOPS because the instruction invalidates the cache line on every write."