Cosmos 3
150. 对英伟达研究副总裁刘洺堉的4小时访谈:Cosmos 3、世界模型、武术、黄仁勋影响我的,和你不需要击败所有对手 adds a first-person builder account through Liu Ming-Yu / 刘洺堉. Liu describes Cosmos 3 as the point where Cosmos Lab moved from several separate predict, transfer, reason, and policy directions into one Omni-Model for world foundation model use.
In this source, Cosmos 3 combines language, video, audio, and action because Physical AI agents observe the world, communicate with people, and change physical state. Liu says the model uses a two-tower design for discrete and continuous signals, making post-training easier for customers with different robot, vehicle, or simulation needs.
Cosmos 3 is discussed in 170: 【具身季报 26Q2】世界模型大风不停,和不想被贴标签的人 as Nvidia’s Q2 2026 marker for productized World Models. Chen Zhe Peter describes it as a more open, product-level omni-world-model release with Super, Nano, and Edge variants for different deployment contexts.
The episode says Cosmos 3 can handle and generate multiple modalities, including text, images, video, sound, and action, and uses a Mixer of Transformers architecture combining autoregressive and diffusion transformer components. Its significance in the wiki is not only model capability, but the fact that a major infrastructure company can release a world-model stack that supports World Action Models, robot simulation, and World Model VLA Fusion without needing the model itself to be the whole robot policy.
Connections
- Nvidia — company associated with the product in the episode.
- Liu Ming-Yu / 刘洺堉 and Cosmos Lab — guest and internal team behind the source’s builder account.
- World Models, World Action Models, and World Model VLA Fusion — model-route concepts Cosmos 3 extends.
- World Foundation Models, Robot Data Scale Up, and Robot Generalization Performance Tradeoff — source-specific frame for Cosmos as reusable Physical AI infrastructure.
- Embodied AI, Physical AI, and Robotics Simulation Evaluation — physical-AI context where the model may matter.