Updated · 4 episodes · 4 shows · 4 source notes

entity Topics: Technology

Scale AI

Overview

Scale AI is an AI-data company founded by Alexandr Wang that developed from manual labeling into autonomous-vehicle data, defense work, generative-AI feedback, and agent-era post-training and evaluation services.

Current Profile

The source set shows Scale AI as an example of AI Data Infrastructure changing with model demand. Its early work included manual image and text labeling, followed by multimodal sensor workflows for autonomous vehicles and a government and defense branch. After ChatGPT, Wang says the company rapidly moved staff toward generative-AI data and began emphasizing Agent Data: records of how people gather information, check constraints, decide, and act during real tasks.

The data-industry sources place Scale between a factory and a learning system. 谢晨 uses it as the industrialized Data Factory stage after ImageNet, while newer work adds expert-written tasks, rubrics, environments, verifiers, and research methods for post-training. The latest source presents 何韵中 as a Scale researcher and argues that rubrics and RL environments are complementary: environments provide tools and state, while programmatic checks, rubrics, specialist models, and human judgment evaluate outcomes.

Scale’s potential advantage is therefore not only labeling capacity. The episode argues that suppliers may remain valuable when they can acquire private projects, negotiate software and data rights, recruit experts, match task difficulty to a target model, test for reward hacking, and continually identify new domains. These market and product claims remain source-reported; the sources do not provide audited current revenue, product mix, customer contracts, or comparative training gains.

Key Characteristics

  • Data infrastructure spanning labeling, sensor data, defense imagery, generative-AI feedback, and agent work.
  • Operations-heavy production with human labor, quality control, customer-specific workflows, and research methods.
  • Agent-era focus on tasks, process traces, expert rubrics, execution environments, and verifiers.
  • Procurement capability involving real projects, private data, commercial tools, rights, and specialist access.
  • Training-system work that includes difficulty calibration, successful trajectories, anti-reward-hacking checks, and recipes.
  • Business profile exposed to rapid data-product commoditization and continual demand shifts.

Evidence

Qualifications

The sources mix founder testimony, guest interpretation, and technology-news reporting. Customer lists, contract values, staff allocation, Meta-related claims, current strategy, and commercial forecasts are source-scoped. Neither a large contributor network nor exclusive data guarantees task authenticity, representative coverage, verifier quality, model improvement, or defensible economics.

What Changed

  • Reframed agent-era work as complete task-and-verification systems rather than process traces alone.
  • Added vertical procurement, licensing, expert recruitment, and continual task discovery as possible supplier advantages.
  • Added task calibration and reward-hacking resistance to the post-training profile.

Relationships

Sources

4 source notes across 4 shows
  1. Bytes: Week in Review - Apple's new CEO, Meta's latest AI play, and Roblox's safety updates Marketplace Tech
  2. Alexandr Wang on Scale and AI Data Infrastructure The Social Radars
  3. 134. 【数据的综述】和谢晨聊,新时代的石油、历史、版图、数据金字塔、定价与Recipe 张小珺Jùn|商业访谈录
  4. E253|谁在给大模型出题、卖题、判卷?聊聊AI数据行业的野蛮生长 硅谷101