Updated · 6 episodes · 5 shows · 6 source notes

concept Topics: Technology

Embodied Data Pyramid

Definition

The embodied data pyramid is a robotics data strategy that combines scarce high-fidelity robot or teleoperation data, scalable simulation and structured physical data, human first-person or motion data, and broad internet-scale video or image priors.

Current Synthesis

The wiki’s current synthesis is that no single data layer solves embodied intelligence. Real robot data is most grounded but expensive and narrow; simulation and structured 3D can multiply tasks and evaluation; tactile and sensor data capture contact details; human first-person video and internet video supply broader scene priors. The All-In robotics special adds a 1X variant that puts high-quality teleoperation at the top, then human sensor data, egocentric video, and general video, using human-like embodiment to make lower layers more useful.

Key Claims

  • Real robot and teleoperation data are high-value because they contain action, sensor, body, latency, and failure information.
  • Simulation is the scalable middle layer only when it supports physical consistency, counterfactual action, and useful evaluation.
  • Human first-person, motion-capture, and internet video can provide scene and task priors but may lack the force, contact, and body-specific data needed for control.
  • Tactile sensing creates a special data layer because contact deformation, friction, slip, and force are closer to manipulation ground truth than ordinary video.
  • Structured 3D and sim-to-real methods can fill gaps left by raw video or narrow teleoperation traces.
  • Human-like robots can make human video more transferable, but that does not remove safety, privacy, and control-data limitations.

Evidence

Counterevidence & Qualifications

The bounded sources disagree on weighting. Xie Chen emphasizes simulation and recipes because real robot data is too costly to scale alone; Shen Yujun emphasizes real-machine data and stricter cleaning; 1X emphasizes transfer from human-like video; tactile sources argue that visual data is incomplete without force and contact. The synthesis is a portfolio view, not a settled recipe.

What Changed

  • Added 1X’s explicit teleoperation-to-general-video data hierarchy.
  • Clarified that humanoid embodiment can improve human-video transfer while leaving action, safety, and privacy limits intact.

Sources

6 source notes across 5 shows
  1. E244|端到端vs上下分层:机器人路径之争,正在转向? 硅谷101
  2. 从会跳舞到有感知,触觉是机器人通往智能的门票吗?| S10E19 What's Next|科技早知道
  3. 170: 【具身季报 26Q2】世界模型大风不停,和不想被贴标签的人 晚点聊 LateTalk
  4. 134. 【数据的综述】和谢晨聊,新时代的石油、历史、版图、数据金字塔、定价与Recipe 张小珺Jùn|商业访谈录
  5. 147. 和蚂蚁灵波沈宇军聊:机器人原生基础模型、大脑和本体的关系、预训练与数据scale up、老师汤晓鸥 张小珺Jùn|商业访谈录
  6. The $1/Hour Worker: Four Robotics CEOs on Humanoids at Home, China's Threat, and the End of Dangerous Jobs All-In with Chamath, Jason, Sacks & Friedberg