Updated · 1 episodes · 1 show · 1 source notes

entity Topics: Technology

孙一游 / Sun Yiyou

Overview

孙一游 is presented in the source as a University of California, Berkeley researcher working on agent evaluation and participating in Agent’s Last Exam.

Current Profile

Sun’s contribution is a benchmark-governance and task-authenticity frame. He argues that optimizing for a benchmark is not inherently empty when its tasks represent valuable real work, the held-out evaluation data are not sold as training material, and capability improvements transfer beyond one leaderboard. The benchmark should reveal the gap between current models and a desired future capability, while training data address current weaknesses without collapsing the test into the lesson.

He also describes the difficulty of sourcing professional tasks. Real projects can contain private databases, copyrighted software, sensitive records, or unverifiable work claims; expert submissions therefore need compliance screening, process evidence, and sometimes review by another expert. Project scale, task counts, domain coverage, and future plans remain source-reported snapshots.

Key Characteristics

  • Agent-evaluation researcher associated with Agent’s Last Exam.
  • Treats real-work relevance and held-out integrity as prerequisites for meaningful benchmark optimization.
  • Emphasizes the difference between measuring a future capability and supplying training data for current gaps.
  • Highlights sparse coverage of professional workflows and the cost of expert review.
  • Treats contributor incentives and fabricated expert work as benchmark-quality risks.

Evidence

Qualifications

The source does not supply a full academic biography, publication list, formal project governance policy, or independently verified coverage analysis. The profile records only the episode’s account.

What Changed

  • Added Sun as a source for real-work benchmark design, held-out evaluation integrity, and expert-submission verification.

Relationships

Sources

1 source notes across 1 show
  1. E253|谁在给大模型出题、卖题、判卷?聊聊AI数据行业的野蛮生长 硅谷101