World Model VLA Fusion
150. 对英伟达研究副总裁刘洺堉的4小时访谈:Cosmos 3、世界模型、武术、黄仁勋影响我的,和你不需要击败所有对手 adds the Cosmos 3 implementation view through Liu Ming-Yu / 刘洺堉. Cosmos 3 folds language, video, audio, and action into one model family, using a two-tower design for discrete and continuous signals. The source therefore makes fusion less a market label and more a practical architecture problem for world foundation models that customers can post-train.
World model VLA fusion is the route emphasized in 170: 【具身季报 26Q2】世界模型大风不停,和不想被贴标签的人: World Models and Vision Language Action Models should not be treated as a clean either-or. Chen Zhe Peter argues that VLA models are strong at instruction/action generation, while world models improve state prediction, environment modeling, and future-outcome simulation.
The source uses Cosmos 3, Physical Intelligence’s Pi 0.7, and Generalist Gen 1 to show three variants of the convergence. Cosmos 3 represents a productized omni-world-model stack, Pi 0.7 adds lightweight future-image prediction to a VLA route, and Generalist resists labels by emphasizing direct physical-interaction data.
从会跳舞到有感知,触觉是机器人通往智能的门票吗?| S10E19 adds a touch-modality extension to the same convergence. Eric Li Zhiqiang / 李志强 says VLA and world-model builders are reaching a vision/language ceiling in physical tasks, so Tactile Sensing and Tactile Transformer Encoder may be needed for robot models to reason about force, contact, texture, softness, and slip.
147. 和蚂蚁灵波沈宇军聊:机器人原生基础模型、大脑和本体的关系、预训练与数据scale up、老师汤晓鸥 adds 蚂蚁灵波’s VA-style route. 沈宇军 separates digital-world video models from robot models: robots need real-time, forward-time, action-relevant modeling, and the source says Lingbo’s Video, World, and VA2.0 work feed into that physical-world convergence.
Key Claims
- Robot action policies benefit from predicting how the environment will change, not only from mapping vision and language directly to actions.
- The labels “VLA,” “world model,” and “world action model” may be temporary as robot models absorb video generation, action-conditioned prediction, and policy generation into one architecture.
- The fusion view creates a middle path between Aether AI’s stricter Causal World Models thesis and practical VLA deployment work by companies such as Physical Intelligence.
- A tactile extension would make the fusion multimodal in the physical sense: not just seeing and predicting future images, but sensing contact forces while acting.
- A physical-world VA route differs from ordinary video generation because it has to support immediate action rather than only produce high-quality frames.
- Cosmos 3 adds the product-stack version: model consolidation can reduce user confusion and maintenance cost while making language, sensory prediction, and action available in one post-trainable base.
Connections
- World Models, Vision Language Action Models, and World Action Models — model families being fused.
- Cosmos 3, Physical Intelligence, and Generalist — examples in the source.
- Embodied AI and Physical AI — deployment contexts where the fusion matters.
- Tactile Sensing, Optical Tactile Sensing, TouchNet, and Tactile Transformer Encoder — touch modality, sensor route, dataset, and encoder proposed by the What’s Next source.
- 蚂蚁灵波 / Ant Lingbo, 沈宇军 / Shen Yujun, Embodied Native Foundation Models, and Robot Data Scale Up — robot-native VA and data-scale route added by episode 147.
- Liu Ming-Yu / 刘洺堉, Cosmos Lab, and World Foundation Models — Nvidia builder account added by episode 150.