从蒸馏到合成数据到 RSI,模型竞争的下一个焦点是什么?|对谈 Evolvent AI 联创孟繁青
Summary
This 42章经 episode interviews [[MengFanqing|孟繁青]], co-founder of [[EvolventAI|Evolvent AI]], on how model competition is moving from post-training recipes and data toward [[RecursiveSelfImprovement|RSI]]. The episode argues that post-training work is becoming more service-like inside model labs, making Synthetic Agent Data, Environment-Based Agent Benchmarks, evaluation design, and fast iteration increasingly central. Its larger synthesis is that Model Distillation / 模型蒸馏 can accelerate weaker models, but durable competition still depends on architecture, pretraining, high-quality data, environments, organization, and whether systems can turn feedback into future capability through RSI Data.
Key Claims
- Meng Fanqing / 孟繁青 says large-lab post-training has become more engineeringized: researchers often submit data and training requests while stable internal services handle much of the SFT or RL workflow.
- Environment-Based Agent Benchmarks are replacing static question-answer tests in some areas: the benchmark becomes a simulated work environment where an agent operates tools, receives feedback, and is scored on multiple behaviors.
- Data quality is hard to judge before training; practical validation still depends on whether a model improves without benchmark hacking or leakage.
- Chinese model labs are narrowing the gap with overseas frontier labs, and the source attributes this to constrained-resource architecture and efficiency work as well as data and organization, not only to Model Distillation / 模型蒸馏.
- Kimi and DeepSeek are used as examples of domestic model work shaped by architecture or efficiency pressure, including Kimi’s linear-attention direction and DeepSeek’s Multi Latent Attention.
- Model Distillation / 模型蒸馏 is framed as a useful shortcut and acceleration tool, especially for lagging models with larger learning margin, but not as the decisive reason domestic models have become strong.
- [[RecursiveSelfImprovement|RSI]] is defined broadly as giving a model an environment and a goal, letting it act, observe feedback, modify prior actions, and try to break through its previous ceiling.
- The episode rejects a separate “RSI base model” category: RSI ability is expected to be increasingly internalized by stronger base models through pretraining, world-model-like knowledge, and longer context.
- [[ModelHarnessCoEvolution|Harness]] layers are expected to become simpler but not disappear, because models are trained with the harnesses they will use and vertical applications may still need workflows around smaller models.
- Synthetic Agent Data can overlap with distillation when the exploration agent is an outside stronger model such as Claude, but complex environments, task design, and scoring make this more than copying teacher answers.
- RSI Data is presented as a likely next data demand: trajectories where one model helps train or improve another smaller model, including long-running environment, training, evaluation, and revision traces.
- Personal AI usage traces are unlikely to sell directly because they are noisy, hard to clean, and raise compliance problems.
- The data-company filter is tightening: teams need hands-on researchers who can build environments, run training, clean trajectories, and follow rapidly changing model-lab demand.
- The source treats Auto Research and AI For Science as important next directions after coding pipelines stabilize, while keeping organization speed and technical execution as durable differentiators.
- Multimodal work is not dismissed as outside the intelligence mainline, but the source says text is more compressed and easier to optimize than noisy multimodal inputs.
Key Quotes
“post-training 很多时候是在做数据” — Meng’s summary of where post-training leverage has moved.
“给模型一个环境和目标” — the episode’s compact RSI definition.
“蒸馏不是决定性因素” — the source’s position on domestic model progress.
Connections
- 42章经, Meng Fanqing / 孟繁青, and Evolvent AI — show, guest, and startup context.
- RSIbench-data, RSI Data, Synthetic Agent Data, and Environment-Based Agent Benchmarks — new data and benchmark branch added by this source.
- Recursive Self-Improvement, Auto Research, AI For AI, AI For Science, and ML Coding — broader AI-for-AI and research automation context.
- Agent Post-Training, Model Post-Training Bottleneck, Agent Data, AI Data Infrastructure, and Data Pricing In AI — post-training and data-market context.
- Agent Evaluation Benchmarks, Agent Harness, Model Harness Co-Evolution, AI Verification, and World Models — environment, harness, verifier, and feedback mechanism.
- Model Distillation / 模型蒸馏, On-Policy Distillation, Kimi K3, and Model-Infra Co-Design — distillation and model-factory context.
- Kimi, Moonshot AI / 月之暗面, DeepSeek, Zhipu AI, ByteDance, and Tencent Hunyuan / 腾讯混元 — domestic model-lab comparison set.
- Claude, OpenAI, and Anthropic — outside model and frontier-lab references used in the data and Auto Research discussion.
- AI Organization Design and Research Taste — organization and hands-on research judgment as competitive bottlenecks.
Contradictions
- No direct contradiction found.
- The source reinforces 178: 与田渊栋聊 RSI:模型自进化如何到来? by treating RSI as AI improving AI, but shifts emphasis from research taste and scaling dynamics toward data, environments, harnesses, and long-running trajectories.
- The source qualifies E246|何谓蒸馏?聊聊硅谷如何看中国开放模型逼近前沿 by arguing that distillation is a useful acceleration route rather than the decisive explanation for Chinese model progress.
- The source qualifies AI Training Data Scarcity and Model Collapse concerns by saying high-quality data can keep growing through agent and synthetic loops, while also warning that synthetic data is not infinite and still needs environment, scoring, diversity, and verification.