Updated · 1 episodes · 1 show · 1 source notes
孙一游 / Sun Yiyou
Overview
孙一游 is presented in the source as a University of California, Berkeley researcher working on agent evaluation and participating in Agent’s Last Exam.
Current Profile
Sun’s contribution is a benchmark-governance and task-authenticity frame. He argues that optimizing for a benchmark is not inherently empty when its tasks represent valuable real work, the held-out evaluation data are not sold as training material, and capability improvements transfer beyond one leaderboard. The benchmark should reveal the gap between current models and a desired future capability, while training data address current weaknesses without collapsing the test into the lesson.
He also describes the difficulty of sourcing professional tasks. Real projects can contain private databases, copyrighted software, sensitive records, or unverifiable work claims; expert submissions therefore need compliance screening, process evidence, and sometimes review by another expert. Project scale, task counts, domain coverage, and future plans remain source-reported snapshots.
Key Characteristics
- Agent-evaluation researcher associated with Agent’s Last Exam.
- Treats real-work relevance and held-out integrity as prerequisites for meaningful benchmark optimization.
- Emphasizes the difference between measuring a future capability and supplying training data for current gaps.
- Highlights sparse coverage of professional workflows and the cost of expert review.
- Treats contributor incentives and fabricated expert work as benchmark-quality risks.
Evidence
- Role and project - E253|谁在给大模型出题、卖题、判卷?聊聊AI数据行业的野蛮生长 introduces Sun as a Berkeley agent-evaluation researcher and Agent’s Last Exam participant.
- Benchmark boundary - E253|谁在给大模型出题、卖题、判卷?聊聊AI数据行业的野蛮生长 attributes to him the prohibition on selling evaluation tasks as training data and the defense of real-work-aligned optimization.
- Expert-data quality - E253|谁在给大模型出题、卖题、判卷?聊聊AI数据行业的野蛮生长 describes compliance screening, synthetic submissions, process evidence, and expensive specialist review.
Qualifications
The source does not supply a full academic biography, publication list, formal project governance policy, or independently verified coverage analysis. The profile records only the episode’s account.
What Changed
- Added Sun as a source for real-work benchmark design, held-out evaluation integrity, and expert-submission verification.
Relationships
- Agent’s Last Exam - evaluation project he participates in.
- University of California, Berkeley - research affiliation stated by the source.
- 何韵中 - co-guest discussing post-training data and environments.
- Benchmark–Training Data Separation / 评测集与训练数据隔离 - governance boundary central to his account.
- Human Data Contributor Incentive Alignment / 人类数据贡献者激励对齐 - submission-quality and verification problem raised by the project.
- Agent Evaluation Benchmarks - broader evaluation category his work extends.