Updated · 1 episodes · 1 show · 1 source notes

entity Topics: Technology

Agent’s Last Exam

Overview

Agent’s Last Exam is an agent-evaluation project described in the source as collecting verifiable tasks from many professional and engineering subdomains so models can be tested on work beyond static question answering.

Current Profile

The project is presented as an attempt to transfer the useful properties of coding benchmarks into broader work: a task should matter, run in an appropriate tool environment, and admit credible verification. Experts contribute or reconstruct real projects, while the project screens private, copyrighted, sensitive, fabricated, or otherwise noncompliant material.

Its public benchmark and any adjacent training-data activity must remain separate. 孙一游 argues that publishing or selling the held-out evaluation tasks for training would contaminate the leaderboard, even though benchmark builders may use their domain knowledge to develop distinct training material. The source reports approximately 150 public tasks across more than 50 subdomains at the time of discussion and a plan to exceed 1,000 tasks; these figures are time-bound and not independently checked here.

Key Characteristics

  • Cross-domain benchmark centered on agent completion of professional and engineering work.
  • Task design that combines instructions, tools or software, an execution environment, and verification.
  • Expert-sourcing model using prior projects while screening privacy, copyright, and compliance risks.
  • Held-out benchmark integrity separated from adjacent training-data creation.
  • Quality-control burden involving fabricated work, process evidence, and costly expert review.
  • Source-reported expansion from roughly 150 public tasks toward more than 1,000.

Evidence

Qualifications

The source is a podcast summary, not the project’s technical paper, task repository, or current leaderboard. Task totals, planned scale, domain coverage, acceptance rules, and reported biology-workflow coverage remain source-scoped and time-sensitive. Meaningful work alignment does not by itself eliminate benchmark leakage, narrow specialization, flawed tests, or verifier gaming.

What Changed

  • Added the project as a concrete cross-domain environment benchmark.
  • Established held-out evaluation integrity and contributor verification as core parts of its profile.

Relationships

Sources

1 source notes across 1 show
  1. E253|谁在给大模型出题、卖题、判卷?聊聊AI数据行业的野蛮生长 硅谷101