Blog
Four agent benchmarks fully test none of a customer service representative's seven Core O*NET tasks
We audited the seven Core O*NET tasks of Customer Service Representatives against four agent benchmarks. None of the 28 task-benchmark pairs is fully tested and 8 are partly tested. How to run the same audit on your occupation.
Published September 28, 2026
In our reading of the published task sets, four widely used agent benchmarks, tau-bench, WorkArena, TheAgentCompany, and AppWorld, fully test none of the seven Core ONET tasks of Customer Service Representatives (ONET code 43-4051.00). Of the 28 task-benchmark pairs, eight are partly tested and twenty are not tested at all. Partly tested means the benchmark exercises one operation inside the task or approaches the task from a different role. ONET (the Occupational Information Network) is the US Department of Labor’s database of occupations and the tasks each one involves. A benchmark here means a fixed task set with one scoring rule applied to every model. A Core task is one that at least 67 percent of ONET’s respondents rated relevant to the occupation, with a mean importance of 3.0 or higher. For each pair we asked whether the benchmark makes the agent perform the operations the task requires, change the records the task changes, and grade by reading those records.
Which benchmarks did we audit, and what do they grade?
We audited tau-bench, WorkArena, TheAgentCompany, and AppWorld, four benchmarks whose papers describe customer-facing or office work and whose grading mostly reads system state rather than a transcript.
- tau-bench (τ-bench, June 2024) simulates a customer talking to a retail or airline support agent. It holds 115 retail and 50 airline tasks and grades by comparing the database state at the end of the conversation with an annotated goal state.
- WorkArena (March 2024) runs 33 tasks on the ServiceNow platform: forms, list filtering, knowledge base lookups, service catalog orders. Its successor WorkArena++ extends this to 682 tasks.
- TheAgentCompany (December 2024) places an agent inside a simulated software company with GitLab, OwnCloud, Plane, and RocketChat, and scores 175 tasks with checkpoints, some of them judged by a language model. In the paper’s current version (September 2025), the top agent completed 30 percent; the December 2024 original reported 24 percent.
- AppWorld (July 2024) gives an agent nine consumer apps with 457 APIs and 750 tasks, graded with state-based unit tests.
Each benchmark was built to measure something narrower than one occupation’s full job. The audit asks how much of one occupation’s Core task list the four cover between them.
Two newer benchmarks were built around occupations and end-state grading from the start, which is the direction this audit argues for. Agents’ Last Exam (June 2026), from UC Berkeley’s RDI, organizes over a thousand expert-sourced tasks into 55 sub-fields defined with reference to O*NET and the 2018 SOC, and grades verifiable outcomes. Its hardest tier averaged a full pass rate below 1 percent. Zapier’s AutomationBench (April 2026) runs 600 public tasks across sales, marketing, operations, support, finance, and HR in simulated business systems and grades the end state only: whether the correct data ended up in the right systems. We have not audited either against the Core list. The method below applies to them as it does to the four above.
“We built an AI benchmark that measures real work.”
Zapier's co-founder announcing AutomationBench. The benchmark's grading reads the final state of the business systems, which is the third of the three conditions in the audit below.
What did the audit find?
None of the 28 task-benchmark pairs is fully tested: 8 are partly tested and 20 are not tested.
| Core task (O*NET importance) | tau-bench | WorkArena | TheAgentCompany | AppWorld |
|---|---|---|---|---|
| Confer with customers to provide information, take or enter orders, cancel accounts, or obtain details of complaints (4.67) | partly | partly | not tested | partly |
| Keep records of customer interactions or transactions, including actions taken (4.53) | not tested | not tested | partly | not tested |
| Check to ensure that appropriate changes were made to resolve customers’ problems (4.39) | not tested | not tested | not tested | not tested |
| Contact customers to respond to inquiries or notify them of results or adjustments (4.22) | not tested | not tested | not tested | not tested |
| Determine charges for services, collect deposits or payments, or arrange for billing (4.18) | partly | not tested | not tested | partly |
| Complete contract forms, prepare change of address records, or issue service discontinuance orders (4.10) | not tested | partly | partly | not tested |
| Refer unresolved customer grievances to designated departments for further investigation (4.07) | not tested | not tested | not tested | not tested |
Record-keeping, the second-ranked task, is fully graded by none of the four. The top task (importance 4.67) is conferring with customers to provide information, take or enter orders, cancel accounts, or obtain details of complaints. tau-bench exercises the order and cancellation part of that sentence with a simulated user and a database check. WorkArena reaches it through service catalog orders placed from the employee side, and AppWorld through orders and account actions the agent takes for the customer. The second task (4.53), keeping records of the interaction and the actions taken, is fully graded by none of the four. tau-bench grades the order state and leaves the call log unread. TheAgentCompany produces documents and project records, but inside a software team, so we marked it partly.
Follow-through is absent. Three of the seven Core tasks of Customer Service Representatives come after the request is handled. They are checking that a promised change was made (4.39), contacting customers to respond to inquiries or report results (4.22), and referring an unresolved grievance to another department (4.07). None of tau-bench, WorkArena, TheAgentCompany, or AppWorld exercises any of them. In all four benchmarks, the episode ends once the immediate request is resolved.
Payments and forms appear as single operations. Determining charges and collecting payments (4.18) appears in tau-bench only as changing a pending order’s payment method or entering paid baggage on a booking. In AppWorld it appears as a user sending money through a payment app. Neither grades determining the charge itself. Completing contract forms and service orders (4.10) is partly covered by WorkArena’s form filling and TheAgentCompany’s document tasks, and neither takes place in a customer’s account.
One task outside the grid deserves a sentence. Resolving complaints by exchanging merchandise, refunding money, or adjusting bills is rated 4.19, above four of the seven Core tasks, but O*NET tags it Supplemental for this occupation, so the Core-only rule leaves it out. It is also the closest match to tau-bench retail’s exchange and return tools. A buyer scoping an environment should decide whether to include high-importance Supplemental tasks before running the audit.
How do you run this audit on your own occupation?
Pull the occupation’s Core tasks from O*NET, read each benchmark’s task set, and mark every task-benchmark pair as tests, partly tests, or not tested; the method takes an afternoon and needs no code.
- Pull the occupation’s Core tasks and importance ratings from O*NET OnLine and sort by importance. OnLine shows importance on a 0 to 100 scale; the 1 to 5 values in this post come from the downloadable database.
- For each benchmark you are considering, read the paper’s task description section and, where it is public, the task list itself. Write one sentence on what the agent must change and what the grader reads.
- For each task-benchmark pair, mark one of three: tests, partly tests, not tested. Tests requires all three conditions: the operations the task requires, the records the task changes, and a grader that reads those records.
- Record your reasoning next to each verdict. Another reviewer should be able to check each one from the sentence beside it.
The audit judges what a benchmark tests. Coverage tells you what has been examined; it says nothing about how well any model performs, and a task no benchmark covers is not evidence of a model weakness. A partly verdict covers a wide range of partial coverage, so if a pair could go either way, mark it not tested and write down why. A May 2026 paper, Designing Benchmarks for Knowledge Work, proposes describing any knowledge-work benchmark by four fields: the represented activity, the tested setting, the required work product, and the evaluated result. The three conditions here are a buyer’s version of the last three.
What should RL environment buyers ask a vendor?
Before training on an RL environment, ask the vendor to show the occupation’s Core task list and mark which tasks the environment exercises end to end, including the record-keeping and follow-through tasks in the audit above. Our grading methodology requires that grading read the state the work leaves behind and that the verifier reproduce a practitioner’s recorded run before it ships. We would hold any environment to that standard, including the four benchmarks above if they were sold as training environments. Our guide to reading O*NET covers the Core list and importance ratings, and tribal knowledge lives in the exceptions covers why expert judgment has to be recorded in an environment rather than written down.
RL Supply Atlas is a map of AI and human work: explore the workflows behind a profession and the AI research connected to them, so you can decide what to investigate next. If you want this audit run for a specific occupation against the benchmarks you are already using, request a sample packet.
Request supply