Kimi Linear: A New Efficient Attention Architecture Outperforming Full Attention
Kimi Linear: An Expressive, Efficient Attention Architecture

We introduce Kimi Linear, a hybrid linear attention architecture that outperforms full attention across short and long contexts. At its core, Kimi Delta Attention extends Gated DeltaNet with finer gating, while our bespoke chunkwise algorithm ensures high hardware efficiency. Our 3B parameter model reduces KV cache usage by up to 75% and achieves six times faster decoding, offering a superior drop-in replacement for existing attention systems.
Kimi Linear can be a drop-in replacement for full attention architectures with superior performance and efficiency, including tasks with longer input and output lengths.