Intel 프로세서에서 subnormal 부동소수점 연산이 50배 느린 이유

Subnormal floating-point numbers are expensive on Intel processors

Intel 프로세서에서 subnormal 부동소수점 연산이 50배 느린 이유

IEEE 표준의 subnormal 부동소수점 수는 Intel 프로세서에서 일반 수에 비해 곱셈이 약 45~50배, 나눗셈이 18배 느리다. Daniel Lemire의 벤치마크에 따르면 Granite Rapids와 Emerald Rapids에서 subnormal이 입력이든 출력이든 성능 저하가 발생하며, 1%만 섞여도 벡터화된 연산 전체가 느려진다. 반면 AMD Zen 5와 ARM 기반 Graviton 5, Apple M4 Max는 subnormal을 거의 완전한 속도로 처리한다.

Intel 프로세서에서 subnormal 수와 관련된 곱셈은 일반 수에 대한 곱셈보다 약 45~50배 느리다.

이 날의 다른 글

2026-09-18