entity Updated 2026-08-13 Topics: Technology

Cosmos 3

150. 对英伟达研究副总裁刘洺堉的4小时访谈:Cosmos 3、世界模型、武术、黄仁勋影响我的,和你不需要击败所有对手 adds a first-person builder account through Liu Ming-Yu / 刘洺堉. Liu describes Cosmos 3 as the point where Cosmos Lab moved from several separate predict, transfer, reason, and policy directions into one Omni-Model for world foundation model use.

In this source, Cosmos 3 combines language, video, audio, and action because Physical AI agents observe the world, communicate with people, and change physical state. Liu says the model uses a two-tower design for discrete and continuous signals, making post-training easier for customers with different robot, vehicle, or simulation needs.

Cosmos 3 is discussed in 170: 【具身季报 26Q2】世界模型大风不停,和不想被贴标签的人 as Nvidia’s Q2 2026 marker for productized World Models. Chen Zhe Peter describes it as a more open, product-level omni-world-model release with Super, Nano, and Edge variants for different deployment contexts.

The episode says Cosmos 3 can handle and generate multiple modalities, including text, images, video, sound, and action, and uses a Mixer of Transformers architecture combining autoregressive and diffusion transformer components. Its significance in the wiki is not only model capability, but the fact that a major infrastructure company can release a world-model stack that supports World Action Models, robot simulation, and World Model VLA Fusion without needing the model itself to be the whole robot policy.

Connections