Blog
How to read O*NET's 18,796 task statements before you commission an RL environment
The US Department of Labor's O*NET database lists 18,796 task statements across 923 occupations, each rated for importance. How to read it when deciding what to benchmark or build.
Published September 24, 2026
The US Department of Labor already publishes a list of the tasks in each occupation. The O*NET database (the Occupational Information Network), in release 30.3 of May 2026, the version our numbers come from, holds 18,796 task statements for 923 occupations, with an importance rating on 17,951 of them. It is free under a CC BY 4.0 license. GDPval, the Anthropic Economic Index, and Agents’ Last Exam all anchor their tasks to it, and it is the full list of work we measure coverage against when we decide which reinforcement learning (RL) environments to build. In this post, an RL environment means a resettable copy of real software plus a grader, and a benchmark means a fixed set of tasks scored the same way for every model. Both need a list of what the work is before anything is built, and O*NET is the public version of that list.
What does O*NET contain?
ONET is a survey-based database of occupations maintained for the Department of Labor. Each occupation carries a code in the ONET-SOC system, an extension of the federal Standard Occupational Classification (Customer Service Representatives are 43-4051.00), a description, and a list of task statements written at the level of detail of a job description. “Keep records of customer interactions or transactions, recording details of inquiries, complaints, or comments, as well as actions taken” is one task.
Release 30.3 has four parts that matter for anyone choosing what to measure:
- Core versus Supplemental. 13,643 tasks are tagged Core, meaning at least 67 percent of respondents rated the task relevant to the occupation and its mean importance was 3.0 or higher. 4,308 are Supplemental and 845 carry no tag.
- Task statements. 18,796 across 923 occupations in the Task Statements table. Among the 893 occupations with at least one Core task, the median has 15 and the range runs from 1 to 36.
- Importance ratings. 17,951 tasks are rated by incumbents (workers in the occupation at the time of the survey) or by occupational experts on a scale from 1 (not important) to 5 (extremely important) in the Task Ratings table.
- Work activities. Every task links to one or more of 2,087 Detailed Work Activities (DWAs), which roll up to 332 Intermediate Work Activities and 41 Generalized Work Activities such as Documenting/Recording Information or Handling and Moving Objects.
O*NET publishes a new release several times a year. Release 31.0, dated August 2026, became the production database after our analysis and lists 18,838 task statements; the structure described here is unchanged.
Which tasks matter most, and how do you tell?
Within one occupation, the tasks that matter most are its Core tasks with the highest importance ratings: filter to Core, then sort by importance. Incumbents rate nearly every task on their own list as important, so the rating is most useful for ranking tasks inside one occupation.
For Customer Service Representatives, the top-rated task (4.67) is conferring with customers to provide information, take or enter orders, cancel accounts, or obtain details of complaints. The second (4.53) is keeping records of those interactions. An environment for that occupation that skips either one leaves out the two tasks incumbents rate highest.
Which AI research is built on O*NET?
GDPval from OpenAI, the Anthropic Economic Index, Agents’ Last Exam from UC Berkeley, and a March 2026 study of 43 agent benchmarks all map their work onto O*NET. A buyer can treat it as a reference that labs and university groups already share.
- GDPval, published in October 2025, covers 1,320 tasks across 44 occupations and was sourced to cover the majority of O*NET work activities for each occupation. A gold subset of 220 tasks is public.
- The Anthropic Economic Index, February 2025, mapped over four million Claude conversations onto O*NET tasks and found 57 percent of usage looked like augmentation and 43 percent like automation.
- How Well Does Agent Development Reflect Real-World Work?, March 2026, mapped 43 benchmarks and 72,342 tasks onto all 1,016 O*NET occupations and found agent development concentrated on programming while human labor and economic value concentrate elsewhere. It proposes three principles for benchmark design: coverage, realism, and granular evaluation.
- Agents’ Last Exam, June 2026, organizes over a thousand expert-sourced tasks into 55 sub-fields and 13 industry clusters defined with reference to O*NET and the 2018 SOC, with verifiable outcomes. Its hardest tier averaged a full pass rate below 1 percent across the harnesses tested.
O*NET is public and maintained, and its task lists were written without reference to any model’s abilities.
Where does O*NET stop being useful?
ONET has four limits for anyone building an RL environment. First, a task statement names the outcome a worker is responsible for and leaves out the steps. The Customer Service Representative task of keeping records of customer interactions says what the representative must record. It does not say that the work happens in a ticketing system, starts with a case being created, and needs a disposition code before the case can close. An environment has to grade those operations, so somebody has to break the statement into them, and that needs a recording of a practitioner doing the work, which our post tribal knowledge lives in the exceptions covers. Second, ONET rates how often a task recurs only in broad categories, and it does not say how long a task takes or which software a given task runs in. Third, it does not say which tasks a model can already do. Use importance to decide which tasks to check first, and measure separately whether a model already does them. Fourth, it is organized by occupation, and work that is sold by the project does not sort cleanly into it. The Remote Labor Index, 240 freelance projects from Scale AI and the Center for AI Safety, found O*NET unsuitable for classifying its projects and estimating coverage, and used the Upwork taxonomy instead.
How RL Supply uses O*NET to scope environments
Our environment catalog is scoped to task families drawn from an occupation’s Core list. Our grading methodology requires that the grader read what the software’s records hold after the task (its state), calibrated against a working practitioner’s recorded run. O*NET supplies the first step of that chain: the published list of what the occupation does. RL Supply Atlas is a map of AI and human work: explore the workflows behind a profession and the AI research connected to them, so you can decide what to investigate next. The waitlist is open. If you already know the occupation you need to train on, request a sample packet and we will scope the task family against its Core list.
Request supply