Blog
Six benchmarks that measure real work grade it in five different ways
GDPval, the Remote Labor Index, JobBench, Agents' Last Exam, AutomationBench, and tau-bench all claim to measure real work. They grade it five different ways, and the grading decides what a score can tell you about training.
Published October 4, 2026
Six public benchmarks describe themselves as measuring real or economically valuable work, and they grade it five different ways: expert pairwise preference, human acceptance, rubric criteria, verified deliverables, and system state. GDPval has experts compare a model’s deliverable with a human’s, blind and pairwise. The Remote Labor Index asks human evaluators whether a deliverable matches the freelancer’s. JobBench scores outputs against rubric chains averaging 35.6 binary criteria per task. Agents’ Last Exam asks for verifiable outcomes checked against expert-provided references. AutomationBench and tau-bench read the state of the system after the episode (one attempt at the task) and compare it with a goal. A benchmark is a fixed task set with one scoring rule applied to every model, and the scoring rule decides what a score means for training. Two of the five methods can serve directly as the reward in a reinforcement learning (RL) environment, and a third can when a program rather than a model applies it.
What frame makes the six comparable?
Hua et al. propose four fields for describing a knowledge-work benchmark in Designing Benchmarks for Knowledge Work (May 2026), the same frame our customer service audit used: the represented activity, the tested setting, the required work product, and the evaluated result. The first two say what work is being stood in for and under what conditions. The last two say what the agent has to leave behind and which part of it the grader reads. Hua et al. note that benchmark papers usually describe themselves by tasks, environments, and metrics, leaving all four of these choices implicit. The fourth field, the evaluated result, splits the six benchmarks here into five grading methods.
How does each one grade?
Each benchmark reads a different evaluated result: GDPval a preference, the Remote Labor Index an acceptance judgment, JobBench a criteria count, Agents’ Last Exam a verified deliverable, and AutomationBench and tau-bench system state.
| Benchmark | Publisher, date | Tasks | Who or what grades | Evaluated result |
|---|---|---|---|---|
| GDPval | OpenAI, October 2025 | 1,320 across 44 occupations | Professional experts compare the model’s deliverable with a human expert’s, blind and pairwise, on the 220-task gold subset; an automated grader is reported with its agreement to the humans | Preference |
| Remote Labor Index | Scale AI and the Center for AI Safety, October 2025 | 240 freelance projects | Human evaluators judge whether the deliverable completes the project at least as well as the human’s; the share that pass is the automation rate, 2.5 percent for the best agent in the paper | Acceptance judgment |
| JobBench | University of Washington and collaborators, May 2026 | 130 across 35 occupations | A language-model judge applies a fact-anchored chain of rubrics, averaging 35.6 binary criteria per task; the strongest of 36 configurations scored 45.9 percent overall | Criteria count |
| Agents’ Last Exam | UC Berkeley, June 2026 | Over a thousand in 55 sub-fields | Each task should admit deterministic checking or an unambiguous rubric tied to observable artifacts, checked against expert-provided references; the strongest case compares a deterministic deliverable directly against a reference output | Verified deliverable |
| AutomationBench | Zapier, April 2026 | Cross-application workflows in six domains | Programmatic assertions against the simulated world state, end state only: whether the correct data ended up in the right systems | System state |
| tau-bench | June 2024 | 165 retail and airline support tasks | The database state at the end of the conversation is compared with an annotated goal state, plus required outputs in the agent’s replies | System state |
“I'm really excited to share our latest: Agents' Last Exam (ALE).”
Dawn Song announcing Agents' Last Exam on June 9, 2026.
Which grading methods carry over into a training reward?
The methods a program can compute against fixed checks, without a person grading each episode, carry over into a training reward. A reward has to be computed thousands of times and has to hold up against an agent that learns to satisfy the grader rather than do the work. System-state grading (AutomationBench, tau-bench) and verified deliverables (Agents’ Last Exam) meet both conditions when the checks are complete. Rubric criteria (JobBench) carry over only partly: its authors anchor each binary criterion to facts in the reference files, but a language-model judge applies them, so the reward still runs through a model’s reading.
Expert pairwise preference (GDPval) and human acceptance (the Remote Labor Index) need a person to grade each episode, so they serve as evaluations and cannot serve directly as training rewards. Using them in a training loop means putting a learned judge or reward model in place of the people. Artificial Analysis’s independent GDPval-AA leaderboard does exactly that, substituting a language-model judge for the human experts, and the substitution reintroduces the failure our post on TypeSafe’s Jev describes: text that persuades the judge without completing the work. Our post on why pinned environments outlast preference datasets covers the cost side of the same split.
What to ask when a vendor cites a benchmark score
When a vendor cites a benchmark result, ask what the evaluated result was: a preference, an acceptance judgment, a criteria count, a verified deliverable, or system state. Only the last two can become a reward directly, and rubric criteria can when a program applies them and an agent cannot pass them without doing the work. Our grading methodology reads the state the task leaves in the software and checks it against written criteria calibrated on a practitioner’s recorded run.
RL Supply Atlas is a map of AI and human work: explore the workflows behind a profession and the AI research connected to them, so you can decide what to investigate next. To see how we grade a sample episode, request a sample packet.
Request supply