Blog
Which O*NET tasks to build into an RL environment first: Core, Supplemental, and importance
In 625 of the 763 O*NET occupations with both tags, a Supplemental task outranks the lowest-rated Core task on importance. A rule for choosing tasks that uses both tags, and how three public benchmarks choose instead.
Published September 30, 2026
Across the 763 occupations in ONET release 30.3 that carry both Core and Supplemental tags, 625 (82 percent) have at least one Supplemental task rated more important than their lowest-rated Core task. A task list built from the Core tag alone therefore leaves out work incumbents rate as important. ONET (the US Department of Labor’s Occupational Information Network) calls a task Core when at least 67 percent of incumbents, the workers in the occupation, rated it relevant with a mean importance of 3.0 or higher. Every other rated task is Supplemental. In 304 of those occupations (40 percent) a Supplemental task outranks the median Core task, and in 58 (8 percent) the single most important task in the occupation is Supplemental. For Payroll and Timekeeping Clerks, filing payroll tax returns is rated 4.69, tied for third of 21 tasks, and is Supplemental. When you scope a reinforcement learning (RL) environment, start with the Core tasks sorted by importance, then check the Supplemental list before you stop.
Why do the tags and the ratings disagree?
O*NET’s Core tag and its importance ratings disagree because the tag depends on how many incumbents perform a task, while the rating measures how much the task matters to those who do. A task that a large share of incumbents do not perform at all can still be rated extremely important by those who do, and it lands in Supplemental. Payroll tax filing is an example: not every payroll clerk files, but those who do rate it close to issuing paychecks (4.69 against 4.72).
Core tasks still rate higher on average. Across all 894 rated occupations, the median Core task is rated 4.13 and 62 percent of Core tasks score 4.0 or higher; the median Supplemental task is rated 3.73 and 30 percent score 4.0 or higher. Among the 763 occupations with both tags, 2,540 of 4,294 Supplemental tasks (59 percent) outrank at least one Core task in their own occupation. Use the Core tag as a first filter, then add the Supplemental tasks the rule below picks up.
How wide is the spread inside one occupation?
Among the 763 occupations with both tags, the median gap between the most and least important Core task is 1.01 points on O*NET’s 5-point importance scale, and in 51 percent of occupations it is a full point or more. The top of the list is usually a cluster: the median occupation has 4 Core tasks within 0.25 of its highest rating.
Which O*NET tasks should an RL environment cover first?
The rule we apply when scoping an environment has three steps:
- Take the Core tasks and sort by importance.
- Add any Supplemental task rated above the occupation’s median Core task.
- Build the top cluster first: every task within 0.25 of the highest rating, then work down.
For Payroll and Timekeeping Clerks that cluster holds ten tasks, from verifying attendance and posting to records (4.74) down to recording employee information changes and issuing pay adjustments (both 4.55). One of them, filing payroll tax returns (4.69), is Supplemental. The occupation’s median Core task is rated 4.55, so step 2 picks it up. The state-based grading those record-keeping tasks need is covered in our post on reward design.
How do public benchmarks choose their tasks?
Three public benchmarks each use a different selection rule, and none of them selects by importance.
- GDPval starts from economic weight: the nine sectors contributing most to US GDP, then 44 occupations within them, then tasks written by professionals to cover the majority of each occupation’s O*NET work activities.
- JobBench (May 2026) starts from what workers want handed off: 130 tasks across 35 occupations chosen because experts identified them as high priority for delegation. Its authors frame this as a correction to selecting by GDP value.
- Agents’ Last Exam (June 2026) starts from projects that experts completed, organized into 55 sub-fields referenced to O*NET and the 2018 Standard Occupational Classification (SOC).
GDP weight measures economic value, delegation priority measures what workers would hand off to an agent, expert projects measure what has been done, and importance measures what incumbents say the job depends on. To train a model on an occupation, select tasks by importance, because importance records what incumbents say the job depends on, and include the Supplemental tasks the rule above adds. Our audit of four agent benchmarks against one occupation’s Core tasks shows what a task list misses when selection ignores importance.
What to ask a vendor about Supplemental tasks
Ask for the occupation’s full rated task list, both tags, sorted by importance, and ask which tasks the environment covers. If the answer is a Core-only list, ask about the Supplemental tasks above the median; for 40 percent of the 763 occupations with both tags there is at least one. Our grading methodology requires a practitioner’s recorded run for each task in scope, and a Supplemental task the practitioner does not perform cannot be recorded, which flags it before the build.
RL Supply Atlas is a map of AI and human work: explore the workflows behind a profession and the AI research connected to them, so you can decide what to investigate next. To see a task list scoped for your occupation, request a sample packet.
Request supply