MOPD Post-Training
MOPD post-training is the source-scoped recipe described in 177: 详解Kimi K3:强到冲击Anthropic估值的模型什么样? for merging domain-specialized expert models back into one model. The episode says [[KimiK3|Kimi K3]] first trains nine domain expert models, then uses MOPD to combine their capabilities.
The operational reason is team modularity. Coding agents, general agents, and generalist behavior can require different data, environments, reward strategies, rollout lengths, harnesses, and algorithms. MOPD lets domain teams deliver trained expert models without forcing every post-training recipe and infrastructure component into one shared pipeline.
Key Claims
- Post-training can be organized around domain expert models before a later merge step.
- The recipe reduces coordination cost across teams working on different agent or reasoning domains.
- In this source, distillation is not only model compression; it is also a way to transfer multiple teacher capabilities into one student model.
- MOPD makes Model Harness Co-Evolution organizational: each domain’s harness and reward design can evolve before capabilities are merged.
Connections
- Kimi K3, Model Distillation / 模型蒸馏, On-Policy Distillation, and Agent Post-Training — source post-training context.
- Agent RL, Agent Harness, Model Harness Co-Evolution, and AI Verification — environment and evaluation implications.
- Zeng Zhiyuan / 曾志远 and Zhao Chenyang / 赵晨阳 — guests explaining the workflow.