entity Updated 2026-08-08 Topics: Technology

Kimi Linear

Kimi Linear appears in 177: 详解Kimi K3:强到冲击Anthropic估值的模型什么样? as the smaller-model predecessor whose experiments informed Kimi K3’s hybrid attention design. The source says the three-to-one pattern of KDA layers to Gated MLA layers came mainly from Kimi Linear-scale experiments rather than from a full search at K3’s reported 3T scale.

The page matters because Kimi Linear turns K3 from an isolated release into a scaling case. KDA, NoPE, and long-context efficiency are presented as ideas tested at smaller scale, then carried into a much larger open-weight model with new serving and training challenges.

Connections