Data catalog
Task Datasets & Rubrics
Long-horizon task datasets and expert-weighted rubrics, elicited from the practitioners who do the work.

Catalog
Choose the supply that matches the work.
Conversational Data
Multi-turn data from work that happens through conversation: negotiations, approvals, and coordination with the outcome attached.
02Verifiers & Rubrics for Verticals
Expert-weighted rubrics and deterministic verifiers for vertical AI, portable into the stack you already run.
03Multimodal Data
Screen, voice, document, and video capture of professional work, aligned to one verified episode.
04Off-the-Shelf Datasets
Licensed cuts of tasks, rubrics, and environments you can sample this week and point your agents at.
05Custom Evals
Private benchmarks built to your requirement and baselined against the experts who do the work.
Specification
What ships in the box
Practitioner-authored
Tasks elicited from verified practitioners in the occupations you are training for.
Grading that compiles
Expert-weighted rubrics resolve to code, so a pass means the work actually landed.
Shared or exclusive
License a shared cut, or commission a task family that stays yours alone.
Remarks / catalog
Every cut ships with tasks, rubrics, verifiers, and an expert baseline run. Sample packets and environment access are available on request.
Request supply
Tell us where your model breaks.
We build the environment.
Sample task packets and environment access for evaluation. Tell us the capability you care about and we will send the relevant cut.