Cloudflare runs Kimi and GLM at scale with 8-bit caches and 4-bit weights
Smaller, faster, safer: running Kimi and GLM at scale

Cloudflare's Workers AI serves Moonshot's Kimi K-series and Z.ai's GLM models on GPUs close to users. To fit these large, long-context mixture-of-experts models into memory, they quantize the KV cache to FP8 and compress weights to INT4, doubling context capacity and boosting decode throughput by up to 55% with no accuracy loss. They also add cache integrity checks to protect shared memory, all while keeping overhead under 1%. These optimizations lower costs and increase capacity for customers.
Quantizing the cache adds a small amount of work per token, since the FP8 attention kernel has to convert values as it reads them.