Blog
Why a preference reward cannot grade the record-keeping tasks in O*NET's largest activity family
Grouping O*NET's 17,951 rated tasks by work activity puts Work Output first at 35 percent of task-to-activity links, with 1,575 tasks on Documenting/Recording Information. Why those tasks need a state-based verifier.
Published October 2, 2026
Reinforcement learning from human feedback (RLHF) trains a model toward outputs a human labeler prefers. Many tasks in ONET (the US Department of Labor’s Occupational Information Network) have no reader: a record is updated or it is not, and a payment posts or it does not. We grouped the 17,951 importance-rated task statements in ONET release 30.3 by their generalized work activities, counting each task once for every activity it links to. Work Output, the family that covers making, moving, operating, and recording, is the largest at 7,412 of 21,177 task-to-activity links, or 35 percent. Inside it, 1,575 tasks link to Documenting/Recording Information, second in the whole database only to Handling and Moving Objects. Interacting With Others is next at 31 percent. A verifier is the check that decides whether a task’s success conditions were met. Tasks whose result is a changed record need a verifier, because a preference reward judges a description of the work instead of the record.
What does RLHF optimize for?
RLHF trains a model to produce outputs that human labelers rank highly. The recipe was described in OpenAI’s 2022 paper, Training language models to follow instructions with human feedback. Collect demonstrations of desired behavior, collect rankings of model outputs, fit a reward model to the rankings, and optimize the policy against that reward. The objective is the labeler’s judgment of the output.
Rating the output fits when a person is going to read it. It is a weak objective when the output is a database row, because the labeler is judging a description of the work rather than the work. A transcript that reads as if the account was canceled and a canceled account are different things, and a preference reward cannot tell them apart without opening the database. Our post on how TypeSafe’s Jev changes RL training environments covers the related failure when a model is used as the grader: text that persuades the grader without completing the operation.
The two instruments now run side by side in public, from the same evaluator. Artificial Analysis grades its GDPval-AA leaderboard by blind pairwise comparison of model deliverables with a language-model judge, aggregated into Elo ratings. It grades its AutomationBench-AA leaderboard, announced July 6, 2026, by deterministic checks on whether the correct data ended up in the right systems across 657 private tasks, with business-rule guardrails that must not be broken. The first is a preference instrument. The second reads state. UC Berkeley’s Agents’ Last Exam takes the second route as well, grading expert-sourced tasks on verifiable outcomes rather than human or model judgment.
“our independent leaderboard for Zapier's AutomationBench”
The same evaluator that scores GDPval deliverables with a model judge scores AutomationBench by reading the end state of the systems. Both methods are legitimate; they measure different things.
What kind of work do O*NET tasks describe?
By work activity, the largest family in ONET 30.3 is Work Output, at 35 percent of task-to-activity links, followed by Interacting With Others at 31 percent, Mental Processes at 19 percent, and Information Input at 15 percent. ONET links every task statement to one or more of 2,087 Detailed Work Activities (DWAs), which roll up through 332 Intermediate Work Activities to 41 Generalized Work Activities. Those 41 sit in the four families defined in the O*NET content model. We joined the Tasks to DWAs table to the activity hierarchy table for the 17,951 rated tasks and counted each task once per generalized work activity it links to, which gives 21,177 task-to-activity links.
The largest activities in each family, by the number of rated tasks linked to them:
- Work Output. Handling and Moving Objects (2,511), Documenting/Recording Information (1,575), Performing General Physical Activities (1,210), Controlling Machines and Processes (700).
- Interacting With Others. Guiding, Directing, and Motivating Subordinates (1,159), Communicating with Supervisors, Peers, or Subordinates (782), Providing Consultation and Advice to Others (745), Assisting and Caring for Others (678).
- Mental Processes. Thinking Creatively (1,133), Analyzing Data or Information (641), Making Decisions and Solving Problems (641), Judging the Qualities of Objects, Services, or People (519).
- Information Input. Getting Information (1,036), Inspecting Equipment, Structures, or Materials (937), Monitoring Processes, Materials, or Surroundings (764).
A large share of Work Output is physical and out of scope for any software environment. The digital portion of Work Output, with much of Information Input and the record-producing parts of Mental Processes, describes work whose result is a change in the software that holds the official record, such as a ticketing system. For those tasks, a grader can check the result by reading that record after the episode.
Which O*NET tasks need which kind of reward?
Tasks whose result is a changed record need a verifier that reads the system after the episode. Tasks whose result is a person’s understanding need an expert’s recorded run and written criteria. Tasks that do both need a reward that covers each half. Group an occupation’s tasks by what the result is, and the instrument follows:
| What the task leaves behind | Examples from Customer Service Representatives (O*NET importance) | Reward instrument |
|---|---|---|
| A changed record | Keep records of interactions and actions taken (4.53); determine charges and collect payments (4.18); complete contract forms and service orders (4.10); check that the change was made (4.39) | The state of the record after the episode: present, correct, consistent with the inputs. Requires real software with a known version so the state is real and resettable, which is how our environment catalog is built. |
| A person’s understanding or decision | Contact customers to respond to inquiries or report results (4.22) | An expert’s recorded run as the reference, with written criteria for an acceptable outcome. A preference rating from a generic labeler is weaker, because the labeler lacks the expert’s standard. |
| Both | Confer with customers to inform, take orders, cancel, or obtain details of complaints (4.67); refer unresolved grievances (4.07) | Both instruments. A reward that grades only the conversation gives no signal for the record-keeping half. |
The first row is the case our grading methodology is built for. The second row is the design we described in tribal knowledge lives in the exceptions: record the expert doing the work, then grade attempts against that record and written criteria. The top two Customer Service Representative Core tasks form a pair from the third row: confer with the customer, then record the interaction and the action taken.
What should buyers ask about an RL environment’s reward?
Ask what the reward reads. If the answer is the opinion of a judge model, a second model asked to rate the transcript, the environment is measuring assistance. The O*NET task it claims to cover is then probably a task whose result is a record that the grader never opens. If the answer is the state of the software after the episode, ask how that grader was calibrated and against whose recorded work. Our posts on how to read O*NET and the customer service benchmark audit show the task lists this applies to.
RL Supply Atlas is a map of AI and human work: explore the workflows behind a profession and the AI research connected to them, so you can decide what to investigate next. To see a sample episode with the state read by the grader, request a sample packet.
Request supply