concept Updated 2026-08-06 Tags: Ai, Model-Training, Post-Training

Model Post-Training Bottleneck

Model post-training bottleneck is the episode’s reminder that a useful large model is not finished when pretraining succeeds. In Vol.114 AI的2025和DeepSeek们的未来 | 对谈复旦张奇教授, [[ZhangQi|张奇]] argues that pretraining supplies much of a model’s knowledge, but post-training decides whether that latent knowledge becomes reliable behavior for a task.

The bottleneck is partly data matching. Zhang says that if a model has already remembered the relevant knowledge during pretraining, a small amount of aligned training data may unlock usable answers; if not, adding supervised data later can fail or even disturb other behavior. The source therefore treats post-training as a costly, expert-heavy search process rather than a simple “add labels” phase.

Key Claims

  • Post-training must match the knowledge and behavior the base model can actually support.
  • More high-quality-looking data is not automatically better if it does not align with what the model already learned.
  • Frontier labs’ advantage may include tacit formulas, evaluation know-how, expert labeling, and large-scale trial-and-error after pretraining.
  • DeepSeek can reduce the visible cost story around pretraining and inference, while the post-training layer can remain expensive and hard to copy.
  • Agent systems intensify the bottleneck because the model must learn reflection, tool use, memory, failure recovery, and environment feedback.

Connections