178: 与田渊栋聊 RSI:模型自进化如何到来?
Summary
This LateTalk episode interviews [[TianYuandong|田渊栋]] on [[RecursiveSelfImprovement|recursive self-improvement]] and the company [[Recursive|Recursive Superintelligence]]. The source argues that RSI is broader than a stronger coding agent: code execution is a necessary early substrate, but open-ended AI research still needs Research Taste, abstraction, direction selection, and AI Verification. Its larger synthesis is that AI may compress research feedback loops and produce useful low-level self-improvement before full automation, while the long-term race depends on whether intelligence improves smoothly through Frontier Model Scaling or through platform-and-breakthrough cycles.
Key Claims
- [[Recursive|Recursive Superintelligence]] is presented as a company trying to use AI to improve AI by finding new models, paradigms, training logic, optimization methods, and research workflows.
- Tian Yuandong / 田渊栋 says model capability and AI coding ability make RSI newly plausible because models can now help discover model weaknesses and execute parts of researcher work.
- The source treats AI Research Feedback Compression as an early RSI signal: AI can turn ideas into code, experiments, and results in minutes or hours rather than days or weeks.
- RSI is not equated with full automation. Humans may remain in the loop, but their work shifts upward toward taste, framing, direction, and deciding which results matter.
- The episode distinguishes current AI safety or coding-agent demonstrations from full RSI: a system may accelerate research without yet improving the next model loop.
- ML Coding is a strong early path because experiment code, benchmarks, kernels, and training scripts can be executed and checked, but breakthrough AI research often lacks a known path.
- The source argues that lower-order RSI already has practical value in algorithm and kernel optimization, while higher-order RSI would require models to derive deep insight from sparse evidence.
- Tian Yuandong / 田渊栋 is skeptical that Frontier Model Scaling alone explains future progress: he accepts scaling as useful, but expects compute, data, and energy limits to require better methods.
- The episode frames the strong-get-stronger question through S-curve dynamics: if intelligence improvement has plateaus and breakthroughs, startups may still find room against large frontier labs.
- Mechanistic Interpretability is treated as both safety infrastructure and discovery infrastructure because better internal understanding could surface insights and improve model design.
- [[Recursive|Recursive Superintelligence]]’s reported early results on NanoChat, NanoChat speed run, and operator optimization are presented as evidence of one general system working across multiple small, verifiable tasks, not proof of open-ended superintelligence.
- The source ties AI Organization Design to RSI: small, hands-on model teams may preserve faster feedback and better technical judgment than layered manager/executor structures.
Key Quotes
“Scaling Law 不是全部” — the episode’s scaling-boundary frame.
“递归不等于完全自动化” — the source’s distinction between self-improvement loops and full autonomy.
“AI 变成科学” — the long-term interpretability and principle-finding aspiration.
Connections
- LateTalk and Tian Yuandong / 田渊栋 — show context and interviewee.
- [[Recursive|Recursive Superintelligence]], Recursive Self-Improvement, Auto Research, and AI For AI — central company and technical frame.
- AI Research Feedback Compression, ML Coding, Model Harness Co-Evolution, and AI Verification — mechanism by which AI can accelerate AI research while still needing checks.
- Research Taste, Human Taste as AI Training Signal / 人的品味作为AI训练信号, Problem Definition In Research, and Human Judgment Under AI — human standards that remain bottlenecks as execution becomes faster.
- Frontier Model Scaling, Mechanistic Interpretability, AI For Science, and Discovery Model — broader question of whether new principles and discovery systems can move beyond scaling alone.
- Anthropic, OpenAI, Claude Code, and Kimi K3 — frontier-lab, coding-agent, and recent LateTalk comparison context.
- Google, FAIR, and AlphaGo — background institutions and systems experience behind Tian’s research judgment.
- AI Organization Design — small-team and hands-on research-organization implications.
Contradictions
- No direct contradiction found.
- The source reinforces 171: 【AI季报 26Q2】从 coding 到 RSI,强者愈强的未来? by treating coding and verifiable benchmarks as early RSI surfaces, but qualifies it by saying coding-agent ability is only a necessary condition, not the full research-intelligence problem.
- The source complements 149. 亲历中美 New Labs 资本狂潮,和清华刘子鸣聊:AI for AI、机制可解释性和 Max Tegmark by adding Tian’s more founder/operator view of AI-for-AI: the key split is not only diligent versus smarter research automation, but also whether future model progress follows smooth scaling or punctuated breakthroughs.