Kimi Linear
Kimi Linear appears in 177: 详解Kimi K3:强到冲击Anthropic估值的模型什么样? as the smaller-model predecessor whose experiments informed [[KimiK3|Kimi K3]]’s hybrid attention design. The source says the three-to-one pattern of KDA layers to Gated MLA layers came mainly from Kimi Linear-scale experiments rather than from a full search at K3’s reported 3T scale.
The page matters because Kimi Linear turns K3 from an isolated release into a scaling case. [[KimiDeltaAttention|KDA]], [[NoPositionEncoding|NoPE]], and long-context efficiency are presented as ideas tested at smaller scale, then carried into a much larger open-weight model with new serving and training challenges.
Connections
- Kimi, Kimi K3, and Moonshot AI / 月之暗面 — model family and company context.
- Kimi Delta Attention / KDA, NoPE / No Position Encoding, and Model-Infra Co-Design — architecture and scale-up branch.
- AI Inference Cost Structure, Agent Inference Workload, and Inference Acceleration Stack — long-context serving context.