World Action Models
World action models, or WAMs, are discussed in 哪条路线,才能通往「世界模型」的终局?|对话黄碧薇:Aether AI 创始人 as an intermediate route between Video Models and full World Models. Huang Biwei treats WAMs as stronger than pure Vision Language Action Models in the short term because abundant video data can help model action-conditioned dynamics.
170: 【具身季报 26Q2】世界模型大风不停,和不想被贴标签的人 adds Nvidia’s three-way taxonomy: Video World Model, Action-Conditioned World Model, and World Action Model. In Chen Zhe Peter’s reading, the embodied-AI community cares most about WAM-like systems because they sit closest to robot policy, but they still fit into a broader World Model VLA Fusion trend rather than replacing every VLA-style system.
171: 【AI季报 26Q2】从 coding 到 RSI,强者愈强的未来? adds the broader AI-quarter frame. Henry Yin says world models became hotter because RL-style world models and video-generation routes began to converge, with action-conditioned prediction as the key bridge between plausible video and robot decision-making.
147. 和蚂蚁灵波沈宇军聊:机器人原生基础模型、大脑和本体的关系、预训练与数据scale up、老师汤晓鸥 adds 蚂蚁灵波’s physical-world version through 沈宇军. The source says digital video generation and robot execution have different requirements: robots need real-time, one-way, action-relevant modeling, so Video, World, and VA-style work only matter if they improve physical action.
173: 对话姚颂:深鉴、东方空间、再出发,「天才少年」十年后 adds Yao Song / 姚颂’s industry-cycle reading. He argues that WAM attention is not only hype because it may push physical-intelligence algorithms forward, but he is skeptical of companies that rely only on an ultimate general model story without a business base, data loop, and Milestone Commercialization path.
Limitation
The episode still does not treat WAM as the end state. Huang argues that a complete route needs Causal World Models: causal variables, causal structures, and transition dynamics grounded in the physical world rather than only action-conditioned video prediction.
Shen adds that action-conditioned modeling also has a data condition: without Robot Data Scale Up, a robot-native VA route can improve fixed tasks but still struggle to generalize to unseen tasks.
Yao adds a company-building limitation: WAM research still has to be embedded in Physical Intelligence System Stack, scenario access, and commercial milestones, or the strongest-financed model-only companies may be the only ones able to survive the long route.
Connections
- Causal World Models — higher-ceiling route in the source.
- Vision Language Action Models — related robot-policy route with a lower ceiling in Huang’s rating.
- Video Models — data and modeling base from which WAM emerges.
- Embodied AI and Aether AI — deployment area and company context.
- Cosmos 3, Nvidia, and World Model VLA Fusion — Q2 2026 product and taxonomy context from the LateTalk source.
- Physical AI, OpenAI, and Anthropic — Q2 AI-quarter context linking world models back to frontier labs and robotics.
- 蚂蚁灵波 / Ant Lingbo, 沈宇军 / Shen Yujun, Embodied Native Foundation Models, and Robot Data Scale Up — physical-world VA route added by episode 147.
- Yao Song / 姚颂, Striding AI / 正行创新, Physical Intelligence System Stack, and Milestone Commercialization — WAM hype and business-base caution added by episode 173.