Kimi Linear: Hybrid Attention Architecture Beats Full Attention
Kimi Linear's KDA module outperforms full attention, cutting KV cache by 75% and boosting throughput 6x at 1M context.
Kimi Team has released Kimi Linear, a hybrid linear attention architecture built around Kimi Delta Attention (KDA), an extension of Gated DeltaNet with finer-grained gating for more effective use of finite-state RNN memory. A specialized chunkwise algorithm using a custom Diagonal-Plus-Low-Rank transition matrix cuts computation substantially while staying closer to the classical delta rule, delivering strong hardware efficiency.
The team pretrained a model with 3B activated and 48B total parameters, layering KDA together with Multi-Head Latent Attention (MLA). Under an identical training recipe, Kimi Linear outperformed full MLA across short-context, long-context, and RL scaling regimes, while cutting KV cache usage by up to 75% and boosting decoding throughput up to 6x at a 1M-token context length.
The result matters because linear attention has long promised efficiency gains but typically traded away quality versus full attention. Kimi Linear is presented as the first architecture to close -- and in fair comparisons exceed -- that gap, positioning it as a drop-in replacement for full attention in production inference. The team has open-sourced the KDA kernel, a vLLM implementation, and both pretrained and instruction-tuned checkpoints, giving engineers a directly usable path to cheaper long-context inference.