Mixture of Experts
177: 详解Kimi K3:强到冲击Anthropic估值的模型什么样? adds Kimi K3 as a large sparse-MoE routing case. The source says K3 routes each token to a small subset of experts from a very large pool, making load balance a training-stability problem. Quantile Balancing is presented as K3’s answer to the quality-versus-balance tradeoff in auxiliary losses and the tuning burden in fixed-step bias updates.
176: 姚顺宇,来到腾讯300天 adds MoE as part of Tencent’s model-scale and organization discussion. The source frames Hunyuan 3 as a roughly 300B MoE model and says WeChat VLM / 微信 VLM had a 258B MoE model for WeChat-internal use, making MoE a marker of both architecture choice and duplicated internal model investment inside Tencent.
Mixture of Experts is the model-architecture pattern discussed in E228|谷歌TPU能撼动英伟达吗?前TPU工程师首次揭秘 through its impact on TPU topology. Henry says earlier TPU generations struggled with MoE because expert routing requires all-to-all communication, while older 2D Torus layouts mostly connected neighboring chips.
The source treats MoE as a practical example of ASIC Workload Prediction Risk. A chip family can be well suited to dense Transformer workloads and then need topology, optical switching, and compiler adaptation when sparse routing becomes important. Later 3D Torus and configurable switching are presented as attempts to make TPU more MoE-friendly.
147. 和蚂蚁灵波沈宇军聊:机器人原生基础模型、大脑和本体的关系、预训练与数据scale up、老师汤晓鸥 adds a model-training engineering case from 蚂蚁灵波. 沈宇军 says the team failed dozens of times over roughly two months before balancing expert activation in its video-model work, making MoE a practical training-stability problem rather than only a sparse-parameter efficiency idea.
148. 对游凯超3小时访谈:开源Infra、和模型Co-design 、“如果vLLM失败,我们会后悔一辈子” adds the inference-engine version through vLLM and DeepSeek. 游凯超 frames MoE as a serving challenge around fine-grained experts, dynamic routing, expert parallelism, and communication, making it a direct case for Model-Infra Co-Design rather than only a model-architecture choice.
Key Claims
- MoE changes the communication pattern as much as the arithmetic pattern.
- All-to-all expert routing can make AI Cluster Networking a first-order model-performance issue.
- Specialized hardware can adapt to architecture shifts, but only if the shift remains close enough to existing workload assumptions.
- MoE therefore links model design to TPU Pod System Optimization, not only to parameter counts or benchmark scores.
- In the Ant Lingbo source, MoE also links architecture design to training convergence and expert-load balancing inside robot-adjacent video modeling.
- In the vLLM source, MoE also links open-model support to inference-engine scheduling, communication, and deployment complexity.
- In the Tencent source, MoE also becomes an organizational cost question because Hunyuan and WeChat can each pursue sparse large models under limited compute.
- In the K3 source, MoE also becomes a routing-statistics problem because extreme sparsity makes expert-load balance central to successful scale-up.
Connections
- TPU, TPU Pod System Optimization, and XLA Compiler — source hardware/software context.
- Transformer Architecture, DeepSeek, and Gemini — model-family context.
- AI Cluster Networking, AI Chip Specialization, and ASIC Workload Prediction Risk — infrastructure and prediction-risk frame.
- 蚂蚁灵波 / Ant Lingbo, 沈宇军 / Shen Yujun, World Action Models, and Embodied Native Foundation Models — embodied-model engineering context added by episode 147.
- vLLM, 游凯超, DeepSeek, and Model-Infra Co-Design — inference-engine co-design context added by episode 148.
- Tencent Hunyuan / 腾讯混元, WeChat VLM / 微信 VLM, Yao Shunyu / 姚顺宇, and Federated AI Organization — Tencent model-organization context added by LateTalk episode 176.
- Kimi K3, Quantile Balancing, Zeng Zhiyuan / 曾志远, and Model-Infra Co-Design — K3 expert-routing branch added by LateTalk episode 177.