DeepSeek's new V4.1-Flash model slashes KV cache to a quarter of its predecessor

DeepSeek v4.1 Flash

DeepSeek's new V4.1-Flash model slashes KV cache to a quarter of its predecessor

DeepSeek has announced DeepSeek-V4.1-Flash, the smallest model in its new architecture family, featuring native visual understanding and an asymmetric Causal Encoder–Decoder design. With 552B total parameters but only 8B active for input and 16B for output, it delivers faster inference and higher throughput. The KV cache requires just one-quarter the HBM and one-eighth the SSD storage of the previous generation, cutting agent costs. It's now live on the DeepSeek API, with off-peak pricing at half the peak rate.

Cache-hit charges often account for a large share of agent costs. Compressing the cache cuts those costs significantly.

More from this day

2026-09-10