concept Updated 2026-08-13 Topics: Technology

World Model VLA Fusion

150. 对英伟达研究副总裁刘洺堉的4小时访谈:Cosmos 3、世界模型、武术、黄仁勋影响我的,和你不需要击败所有对手 adds the Cosmos 3 implementation view through Liu Ming-Yu / 刘洺堉. Cosmos 3 folds language, video, audio, and action into one model family, using a two-tower design for discrete and continuous signals. The source therefore makes fusion less a market label and more a practical architecture problem for world foundation models that customers can post-train.

World model VLA fusion is the route emphasized in 170: 【具身季报 26Q2】世界模型大风不停,和不想被贴标签的人: World Models and Vision Language Action Models should not be treated as a clean either-or. Chen Zhe Peter argues that VLA models are strong at instruction/action generation, while world models improve state prediction, environment modeling, and future-outcome simulation.

The source uses Cosmos 3, Physical Intelligence’s Pi 0.7, and Generalist Gen 1 to show three variants of the convergence. Cosmos 3 represents a productized omni-world-model stack, Pi 0.7 adds lightweight future-image prediction to a VLA route, and Generalist resists labels by emphasizing direct physical-interaction data.

从会跳舞到有感知,触觉是机器人通往智能的门票吗?| S10E19 adds a touch-modality extension to the same convergence. Eric Li Zhiqiang / 李志强 says VLA and world-model builders are reaching a vision/language ceiling in physical tasks, so Tactile Sensing and Tactile Transformer Encoder may be needed for robot models to reason about force, contact, texture, softness, and slip.

147. 和蚂蚁灵波沈宇军聊:机器人原生基础模型、大脑和本体的关系、预训练与数据scale up、老师汤晓鸥 adds 蚂蚁灵波’s VA-style route. 沈宇军 separates digital-world video models from robot models: robots need real-time, forward-time, action-relevant modeling, and the source says Lingbo’s Video, World, and VA2.0 work feed into that physical-world convergence.

Key Claims

  • Robot action policies benefit from predicting how the environment will change, not only from mapping vision and language directly to actions.
  • The labels “VLA,” “world model,” and “world action model” may be temporary as robot models absorb video generation, action-conditioned prediction, and policy generation into one architecture.
  • The fusion view creates a middle path between Aether AI’s stricter Causal World Models thesis and practical VLA deployment work by companies such as Physical Intelligence.
  • A tactile extension would make the fusion multimodal in the physical sense: not just seeing and predicting future images, but sensing contact forces while acting.
  • A physical-world VA route differs from ordinary video generation because it has to support immediate action rather than only produce high-quality frames.
  • Cosmos 3 adds the product-stack version: model consolidation can reduce user confusion and maintenance cost while making language, sensory prediction, and action available in one post-trainable base.

Connections