Updated · 1 episodes · 1 show · 1 source notes
Agent’s Last Exam
Overview
Agent’s Last Exam is an agent-evaluation project described in the source as collecting verifiable tasks from many professional and engineering subdomains so models can be tested on work beyond static question answering.
Current Profile
The project is presented as an attempt to transfer the useful properties of coding benchmarks into broader work: a task should matter, run in an appropriate tool environment, and admit credible verification. Experts contribute or reconstruct real projects, while the project screens private, copyrighted, sensitive, fabricated, or otherwise noncompliant material.
Its public benchmark and any adjacent training-data activity must remain separate. 孙一游 argues that publishing or selling the held-out evaluation tasks for training would contaminate the leaderboard, even though benchmark builders may use their domain knowledge to develop distinct training material. The source reports approximately 150 public tasks across more than 50 subdomains at the time of discussion and a plan to exceed 1,000 tasks; these figures are time-bound and not independently checked here.
Key Characteristics
- Cross-domain benchmark centered on agent completion of professional and engineering work.
- Task design that combines instructions, tools or software, an execution environment, and verification.
- Expert-sourcing model using prior projects while screening privacy, copyright, and compliance risks.
- Held-out benchmark integrity separated from adjacent training-data creation.
- Quality-control burden involving fabricated work, process evidence, and costly expert review.
- Source-reported expansion from roughly 150 public tasks toward more than 1,000.
Evidence
- Purpose and scope - E253|谁在给大模型出题、卖题、判卷?聊聊AI数据行业的野蛮生长 describes the project as extending verifiable evaluation beyond code into more than 50 subdomains.
- Environment structure - E253|谁在给大模型出题、卖题、判卷?聊聊AI数据行业的野蛮生长 links tasks with tools, sandboxes, programmatic checks, rubrics, and human judgment.
- Data governance - E253|谁在给大模型出题、卖题、判卷?聊聊AI数据行业的野蛮生长 preserves the rule against using the evaluation set as training material.
- Contributor quality - E253|谁在给大模型出题、卖题、判卷?聊聊AI数据行业的野蛮生长 reports synthetic or fabricated submissions and the resulting need for process evidence and expert review.
Qualifications
The source is a podcast summary, not the project’s technical paper, task repository, or current leaderboard. Task totals, planned scale, domain coverage, acceptance rules, and reported biology-workflow coverage remain source-scoped and time-sensitive. Meaningful work alignment does not by itself eliminate benchmark leakage, narrow specialization, flawed tests, or verifier gaming.
What Changed
- Added the project as a concrete cross-domain environment benchmark.
- Established held-out evaluation integrity and contributor verification as core parts of its profile.
Relationships
- 孙一游 - participating researcher and source representative.
- Agent Evaluation Benchmarks - broader class of agent capability tests.
- Environment-Based Agent Benchmarks - task, tool, sandbox, and verifier architecture used by the project.
- Benchmark–Training Data Separation / 评测集与训练数据隔离 - governance boundary between measurement and model improvement.
- Expert Rubric Verification / 专家评分标准验证 - method for evaluating outcomes that lack one exact answer.
- Human Data Contributor Incentive Alignment / 人类数据贡献者激励对齐 - quality-control challenge in expert task sourcing.