vLLM v0.28.0: Major Performance Push for Kimi-K3 and DeepSeek V4

vLLM v0.28.0 ships with 584 commits from 270 contributors, focusing on major performance optimizations for Kimi-K3 and DeepSeek V4. Key highlights include Decode Context Parallel (DCP), fused FlashKDA kernels, SiTU activation for MegaMoE, and adaptive speculative token budgets, delivering up to 60% better TTFT. DeepSeek V4 gains sparse MLA end-to-end support, AMD Quark NVFP4, and ROCm enablement. The release also matures Model Runner V2 with E/P/D disaggregation, introduces tiered KV cache offloading to disk, and adds a Rust frontend with gRPC multimodal inference. New defaults raise max_num_batched_tokens to 16384, and several breaking changes are noted, including bitsandbytes moving to a plugin.

Kimi-K3 also now runs on ROCm with the V2 model runner.

More from this day

2026-08-29