entity Updated 2026-08-08 Tags: Ai, Model, Architecture

Kimi Linear

Kimi Linear appears in 177: 详解Kimi K3:强到冲击Anthropic估值的模型什么样? as the smaller-model predecessor whose experiments informed [[KimiK3|Kimi K3]]’s hybrid attention design. The source says the three-to-one pattern of KDA layers to Gated MLA layers came mainly from Kimi Linear-scale experiments rather than from a full search at K3’s reported 3T scale.

The page matters because Kimi Linear turns K3 from an isolated release into a scaling case. [[KimiDeltaAttention|KDA]], [[NoPositionEncoding|NoPE]], and long-context efficiency are presented as ideas tested at smaller scale, then carried into a much larger open-weight model with new serving and training challenges.

Connections