Kimi Linear: A New Efficient Attention Architecture Outperforming Full Attention

Kimi Linear: An Expressive, Efficient Attention Architecture

Kimi Linear: A New Efficient Attention Architecture Outperforming Full Attention

We introduce Kimi Linear, a hybrid linear attention architecture that outperforms full attention across short and long contexts. At its core, Kimi Delta Attention extends Gated DeltaNet with finer gating, while our bespoke chunkwise algorithm ensures high hardware efficiency. Our 3B parameter model reduces KV cache usage by up to 75% and achieves six times faster decoding, offering a superior drop-in replacement for existing attention systems.

Kimi Linear can be a drop-in replacement for full attention architectures with superior performance and efficiency, including tasks with longer input and output lengths.

More from this day

2026-07-28