concept Updated 2026-08-13 Topics: Technology

Embodied Native Foundation Models

150. 对英伟达研究副总裁刘洺堉的4小时访谈:Cosmos 3、世界模型、武术、黄仁勋影响我的,和你不需要击败所有对手 adds a complementary infrastructure-company route. Liu Ming-Yu / 刘洺堉 does not argue that Cosmos 3 replaces robot-native data or body-specific post-training; he frames it as a world foundation model starting point that Physical AI developers can adapt with their own data, tasks, and embodiments.

Embodied native foundation models are 沈宇军’s core thesis in 147. 和蚂蚁灵波沈宇军聊:机器人原生基础模型、大脑和本体的关系、预训练与数据scale up、老师汤晓鸥. The idea is that robot intelligence should not be treated as a language model or digital video model with a mechanical arm attached. It needs a model stack native to sensors, spatial perception, video time, action, embodiment, and real-time physical execution.

The source makes this a practical strategy for 蚂蚁灵波. The company tries to build a robot brain that can run across many bodies, while leaving body form and scene selection open. In Shen’s framing, language models remain useful for understanding instructions, but the embodied model must learn how to turn that instruction into physically successful action.

Key Claims

  • Physical-world models need sensor-native input because real robot cameras, depth sensors, wrist cameras, and noisy data differ from clean academic datasets.
  • Spatial understanding is not the same as semantic recognition: a robot must know whether something can be reached, touched, blocked, or acted on.
  • Video modeling for robots differs from digital-world generation because robot control needs real-time, forward-time, action-relevant prediction rather than slow, bidirectional, quality-first generation.
  • Cross-embodiment brain work requires many configurations, body parts, camera placements, degrees of freedom, and task distributions.
  • Vision Language Action Models, World Models, World Action Models, and VA-style work may converge when the system has to connect perception, future state, and action.
  • Robot Data Scale Up is the limiting condition: embodied-native architecture is not enough without data that scales and remains useful to the model.
  • Better robot brains will reshape sensor and body requirements, so “brain first” still implies body and sensor co-evolution.
  • Nvidia’s Cosmos source adds a platform-starting-point route: open base models can help, but useful embodied performance still depends on post-training, data fit, and task-specific evaluation.

Connections