Agent Post-Training
179: 蒸馏风暴:一场无人公开谈论的技术竞赛 adds Agent Trajectory Distillation to the post-training map. The source says post-training data can now include full agent task traces, not just prompts and answers, but it distinguishes teacher-behavior imitation from using a stronger model only as an evaluator in reinforcement learning.
从蒸馏到合成数据到 RSI,模型竞争的下一个焦点是什么?|对谈 Evolvent AI 联创孟繁青 adds 孟繁青’s “post-training is data” interpretation. The source says that once internal model-lab training services stabilize, researcher leverage shifts toward Environment-Based Agent Benchmarks, Synthetic Agent Data, correctness checks, anti-cheating filters, difficulty control, and whether training on the data actually improves the model.
Agent post-training is Luo Fuli / 罗福莉’s frame in 138. 对罗福莉3.5小时访谈:AI范式已然巨变!OpenClaw、Agent范式很吃后训练、卡的分配、组织平权 for moving SFT, RL, evaluation, and model adaptation from chat behavior toward real agent workflows. The source says agent systems such as Open Claw and Open Cloud expose different requirements: memory, tools, long context, active tasks, cost routing, skills, and multi-step recovery.
E245|藏在大模型背后的新闻人:GPT们的回复是这样写出来的 adds the conversational-content boundary next to the agent boundary. 东尼 / Tony’s voice-agent case shows that model adaptation can target dialogue feel, follow-up depth, and example-conditioned host behavior before the system becomes a full tool-using agent; Content Engineering therefore feeds post-training quality even when the product surface is conversation.
The concept extends Model Harness Co-Evolution. A model trained only for chat may look strong in isolated answers but behave poorly when a framework asks it to plan, call tools, maintain memory, delegate, and verify results. Agent post-training therefore uses simulated user agents, multi-round interaction data, task traces, workflow feedback, and AI Skills to teach the model how to operate inside an Agent Harness.
Vol.114 AI的2025和DeepSeek们的未来 | 对谈复旦张奇教授 adds a broader Model Post-Training Bottleneck frame through 张奇. The episode argues that post-training has to match knowledge already latent in pretraining and that reinforcement-learning or expert-labeling stages can remain costly even when DeepSeek changes the perceived cost of pretraining and inference.
Key Claims
- Post-training becomes more important when the product surface is an agent framework rather than a chatbot.
- Agent data must include environment feedback, tool results, memory updates, task persistence, and failure recovery.
- SFT and RL can be built from user-agent simulations and real framework traces, not only preference comparisons over one-turn answers.
- Different frameworks may require different adaptation because memory shape, tool affordances, channel structure, and cost routing differ.
- Agent post-training makes AI Coding Verification, Long-Horizon AI, and Model Workflow Fit part of model training rather than only deployment evaluation.
- Post-training can be the hidden reproduction barrier when a model’s visible architecture or cost story is easier to discuss than its data recipes, expert labels, evaluation loops, and failure-recovery training.
- The Evolvent AI source adds that agent post-training data may itself become an RSI surface when a model generates, filters, and uses data to improve another model or future behavior.
- LateTalk episode 179 adds that agent post-training can become commercially and legally sensitive when the most useful trajectories come from restricted closed frontier models.
Connections
- Meng Fanqing / 孟繁青, Evolvent AI, Environment-Based Agent Benchmarks, Synthetic Agent Data, and RSI Data — data and environment branch added by the Evolvent AI source.
- Luo Fuli / 罗福莉, Memo VR, and Xiaomi — source speaker, model series, and team context.
- Content Engineering, AI Answer Evaluation, and Voice Interaction — E245’s conversational-quality extension.
- Open Claw, Open Cloud, and Agent Harness — framework layer that changes post-training targets.
- Agent RL, Training Compute Allocation, and Agent-Optimized Model Architecture — infrastructure, compute, and architecture constraints.
- AI Skills, Persistent Agent Memory, and Agent Self-Evolution — reusable workflow and memory signals for training.
- ML Coding, Research Taste, and Model Harness Co-Evolution — research-loop and co-evolution context.
- 张奇, DeepSeek, and Model Post-Training Bottleneck — vol.114’s broader post-training bottleneck and DeepSeek-cost qualification.
- Agent Trajectory Distillation, Model Distillation Evidence, and AI Model Distillation Governance — distillation boundary added by LateTalk episode 179.