concept Updated 2026-08-07 Topics: Technology

Vision Language Action Models

Vision language action models, or VLA models, are robot models discussed in 哪条路线,才能通往「世界模型」的终局?|对话黄碧薇:Aether AI 创始人 as one route toward Embodied AI. The source treats them as useful for connecting perception, instruction, and action, but limited as a final route to World Models.

131. 印奇出任阶跃星辰董事长的访谈:聪明人的诱惑、取舍、超长链路残酷淘汰赛、阶跃函数和超多元方程 mentions VLA as one of StepFun’s model directions alongside base models and all-modal models. In that episode, VLA is less a detailed robotics architecture than a sign that foundation-model companies are moving from language and multimodal models toward terminals and physical interaction.

具身智能的滔天大泡沫中,他已经把机器人送进300个家庭|对话张翼:未来不远创始人/CEO adds Zhang Yi’s operator view from Weilai Buyuan. He treats VLA as still iterating alongside World Models, so F2 Home Robot uses real-home deployment and engineering integration while the modeling routes mature.

132. 对星海图创始人高继扬的3小时访谈:鲶鱼、曾国藩、Waymo与Momenta的两面、一只狼与许华哲的离开 adds Gao Jiyang’s production-robot architecture through Xinghaitu. He describes a dual system where a vision-language model layer decomposes fuzzy instructions and handles logic, while VLA executes physical actions; he rejects putting every ability into one end-to-end model because endpoint compute and latency still matter.

134. 【数据的综述】和谢晨聊,新时代的石油、历史、版图、数据金字塔、定价与Recipe adds 谢晨’s data-demand view. He says large-model teams with VLA groups often have more compute and reinforcement-learning infrastructure than robot startups, but still need Embodied Data Pyramid, Robotics Simulation Evaluation, and Data Recipe Co-Creation to test whether VLA capability is really improving.

170: 【具身季报 26Q2】世界模型大风不停,和不想被贴标签的人 adds a route-convergence update. Chen Zhe Peter says VLA models are strong at instruction and action generation, while World Models improve prediction of future state; Physical Intelligence’s Pi 0.7 is used as an example of VLA augmented with lightweight future-image prediction. The source also uses Generalist to show that some robot-model companies deliberately avoid the VLA label, even when their work remains adjacent to VLA and real interaction data.

从会跳舞到有感知,触觉是机器人通往智能的门票吗?| S10E19 adds a tactile extension. Eric Li Zhiqiang / 李志强 argues that VLA may need to evolve toward a VTLA-style stack where tactile signals are encoded by Tactile Transformer Encoder, aligned with visual features, and used by the robot backbone to reason about force, texture, softness, friction, and slip.

166: 许华哲再次具身创业:不想错过最大的西瓜 adds a critique from Xu Huazhe’s AI Native Robotics route. He does not reject VLA as a useful vocabulary, but argues that a household robot should not be reduced to small model stitching or one-task policies; the action and behavior layer should move toward Unified Robot Models if the goal is Physical AGI.

E244|端到端vs上下分层:机器人路径之争,正在转向? adds Han Zheng / 韩正’s distinction between deployed end-to-end behavior and training-time structure. He argues that a robot policy can run end-to-end at the edge while still learning through Layered Robot Architecture, Structured 3D Robot Data, and Sim2Real rather than pure imitation of human or teleoperation trajectories.

146. 对Physical Intelligence柯丽一鸣4小时访谈:Pi的开源模型研究,机器人的江湖、族谱与主角 adds K’s language-and-action view from Physical Intelligence. K says he became more open to language as a robot interaction layer after using AI agents for work: language can carry context, planning, and reasoning, but current VLA architectures remain early in how they bind language details to action execution.

147. 和蚂蚁灵波沈宇军聊:机器人原生基础模型、大脑和本体的关系、预训练与数据scale up、老师汤晓鸥 adds 蚂蚁灵波’s deployment-driven VLA update through 沈宇军. Shen says VLA matters because industrial landing reveals what data a robot actually needs, but he places it inside a broader Embodied Native Foundation Models stack with depth, video, world/action modeling, and VA-style work rather than treating VLA as the whole robot brain.

173: 对话姚颂:深鉴、东方空间、再出发,「天才少年」十年后 adds Yao Song / 姚颂’s timing critique from Striding AI / 正行创新. He argues that VLA appeared to approach a bottleneck by late 2025, because more data produced smaller gains or even declines, and says some companies were shifting attention toward Robot Reinforcement Learning, World Models, and World Action Models.

Limitation

Huang Biwei argues that VLA generalization is constrained because the action side is a continuous space. Demonstration data can cover many examples, but cannot exhaustively cover all physical states, object conditions, and action variations a robot may encounter.

Han Zheng / 韩正 adds a manipulation-specific limitation: imitation trajectories may reproduce the motion of opening or screwing without understanding the bottle cap, thread, rotation direction, material contact, or future-state change.

K adds a complementary limitation: the “L” in VLA is useful, but the field has not yet deeply explored the fine-grained relation between language, physical context, and action.

Shen adds a data limitation: VLA quality improves when deployment clarifies the needed tasks, configurations, and data quality, but broad new-task generalization still waits on Robot Data Scale Up.

Yao adds a route-timing limitation: if VLA scaling returns flatten before robot data and evaluation mature, companies may need to combine VLA with reinforcement learning, world/action modeling, and scenario-specific deployment rather than wait for VLA alone to become general.

Connections