Kimi Delta Attention / KDA
Kimi Delta Attention / KDA is the linear-attention mechanism discussed in 177: 详解Kimi K3:强到冲击Anthropic估值的模型什么样? as a central part of [[KimiK3|Kimi K3]]’s architecture. The source says K3 replaces most attention layers with KDA while periodically retaining global attention, producing a hybrid design rather than a pure full-attention or pure linear-attention model.
The tradeoff is memory efficiency versus state complexity. KDA compresses long history into recurrent state, reducing the cache and memory movement that make million-token contexts expensive, but that state can be overwritten rather than simply appended like a traditional KV cache. This makes Prefix Caching, speculative sampling, rollback, and serving-engine implementation harder.
The source frames KDA as an engineering validation of linear attention at frontier-adjacent scale. It does not claim fixed-state attention solves long-context forgetting by itself; the episode says K3’s retained global attention and hybrid structure are important for preserving access to older detail.
Key Claims
- KDA makes long-context inference cheaper by keeping much of history in fixed or slowly growing recurrent state.
- K3’s reported pattern is roughly three KDA layers followed by one global-attention layer.
- KDA complicates prefix reuse because recurrent state is mutable rather than append-only.
- Speculative sampling needs special rollback handling when intermediate KDA states have already been updated.
- The mechanism raises the value of Model-Infra Co-Design because model architecture, cache lifecycle, kernels, and serving engine have to line up.
Connections
- Kimi K3, Kimi Linear, and Moonshot AI / 月之暗面 — model family and source context.
- Agent Inference Workload, AI Inference Cost Structure, Prefix Caching, and Inference Acceleration Stack — serving-cost branch.
- NoPE / No Position Encoding, Attention Residues, and Transformer Architecture — adjacent architecture changes.
- Model-Infra Co-Design, Open Source AI Infrastructure, and [[VLLM|vLLM]] — infrastructure implications.