concept Updated 2026-08-13 Tags: Ai, World-Models, Robotics, Physical-Ai, Foundation-Models

World Foundation Models

World foundation models are [[LiuMingyu|Liu Ming-Yu / 刘洺堉]]’s preferred label for Cosmos 3 in 150. 对英伟达研究副总裁刘洺堉的4小时访谈:Cosmos 3、世界模型、武术、黄仁勋影响我的,和你不需要击败所有对手. The term narrows the overbroad World Models label: instead of simply generating plausible video or simulated scenes, the model should provide reusable starting points for developers building Physical AI systems.

In the source, a world foundation model helps through better data, better starting points, and better environments. Better data includes physical-world and egocentric examples; better starting points mean open pretrained models that customers can post-train for their own tasks; better environments point toward simulated worlds where robots and physical agents can practice before costly real-world deployment.

Key Claims

  • A world foundation model is infrastructure, not just a content product; it should help downstream teams train, evaluate, and adapt physical-world agents.
  • Action is a first-class modality because physical agents change the world rather than only describe or observe it.
  • Useful architecture may mix discrete and continuous signal paths so language, video, audio, and action can be combined without forcing every modality into one representation style.
  • Evaluation has to connect benchmarks and arena comparisons to real customer pain because Physical AI failures are often task, body, scene, and cost specific.
  • Open releases can be strategic even for a large infrastructure company when they grow the developer ecosystem and reveal future hardware, serving, and simulation needs.
  • A single model is unlikely to become the next CUDA by itself; the platform effect comes from model, serving, infrastructure, tools, hardware feedback, and customer workflows together.

Connections