Mechanistic Interpretability
178: 与田渊栋聊 RSI:模型自进化如何到来? adds [[TianYuandong|田渊栋]]’s RSI-centered view. He treats interpretability not only as a safety technique, but also as a possible source of research insight: if models become more understandable, researchers and AI systems may find principles that improve future architectures, training methods, and Recursive Self-Improvement loops.
Mechanistic interpretability is the attempt to understand the internal mechanisms by which neural networks produce behavior. In 149. 亲历中美 New Labs 资本狂潮,和清华刘子鸣聊:AI for AI、机制可解释性和 Max Tegmark, Max Tegmark pushes [[LiuZiming|Liu Ziming]]’s group toward the field in late 2022 and early 2023 because large models looked dangerous if their internals remained opaque.
Liu treats the field as close to a biology of AI. It can reveal useful internal structure, but the source also gives a caution: some neuron-level or highly local explanations may fail to survive changes such as random seed variation. That makes Physics Of AI broader than interpretability alone, because it also asks for training dynamics, controlled experiments, and phase-like regularities.
Key Claims
- Interpretability is tied to AI safety when model capability rises faster than internal understanding.
- Visualization can matter because it lets researchers see more than end-to-end prediction quality.
- Mechanistic claims need robustness checks; an explanation that disappears across seeds may be less fundamental.
- Auto Research could accelerate interpretability work if it can structure experiments and compare mechanisms at scale.
- Liu’s route keeps interpretability close to architecture design rather than treating it only as post-hoc explanation.
- Tian’s source adds that interpretability can be part of the capability route as well as the safety route, because insight into mechanisms can guide better model design and faster discovery.
Connections
- Max Tegmark and [[LiuZiming|Liu Ziming]] — source people.
- Physics Of AI — broader scientific frame in which mechanistic interpretability sits.
- AI Interpretability By AI — adjacent wiki concept about AI assisting interpretability.
- [[KolmogorovArnoldNetworks|KAN]] — source case where internal visualization helps evaluate an architecture.
- AI Alignment Governance and Frontier Model Release Governance — broader wiki safety context.
- Tian Yuandong / 田渊栋, Recursive Self-Improvement, AI For AI, and Discovery Model — RSI branch where interpretability becomes discovery infrastructure.