沈宇军 / Shen Yujun
Shen Yujun / 沈宇军 is the chief scientist of [[AntLingbo|蚂蚁灵波]] in 147. 和蚂蚁灵波沈宇军聊:机器人原生基础模型、大脑和本体的关系、预训练与数据scale up、老师汤晓鸥. The source presents his path as a move from [[TsinghuaUniversity|Tsinghua]], early SenseTime exposure, [[TangXiaoou|Tang Xiao’ou]]’s [[ChineseUniversityOfHongKong|CUHK]] research environment, image generation, and ByteDance visual applications into Embodied AI.
His robotics view is brain-first but not hardware-blind. Shen argues that robot bodies may remain diverse across households, factories, and service scenes, so a reusable robot brain has to handle many embodiments. At the same time, he says brains, bodies, sensors, hands, and data will co-evolve because better models will ask different things of cameras, tactile sensors, latency, and physical form.
Source Position
- Shen treats GAN and image-generation work as useful technical background, but says GAN scale-up became less compute-efficient than newer routes for complex image and video generation.
- He moved toward robotics because physical-world robots make vision, space, action, and sensor grounding central rather than peripheral.
- He argues that Embodied Native Foundation Models should be designed around real sensors, spatial intelligence, video time series, action, and real-time execution.
- He frames Robot Data Scale Up as the missing condition for a robot GPT-1 moment.
- He distinguishes intent execution from intent origination: the current target is completing a given instruction reliably, not autonomous high-level desire or planning.
- He sees large language models as useful for instruction understanding, while the embodied model should learn how to make the action work.
Connections
- [[AntLingbo|蚂蚁灵波]] — company where he leads the robot-brain effort.
- [[AntGroup|蚂蚁集团]] — parent-company context for the embodied-AI bet.
- [[TangXiaoou|汤晓鸥]], SenseTime, Tsinghua University / 清华大学, 香港中文大学 / Chinese University of Hong Kong, and ByteDance — career and research-lineage context.
- Embodied Native Foundation Models, AI Native Robotics, Vision Language Action Models, and World Model VLA Fusion — model-route concepts tied to his interview.
- Robot Data Scale Up, Real Robot Data Strategy, and Embodied Robot Data Paradigms — data bottlenecks he emphasizes.