Method
Methodology
How RLSupply turns expert work into environments and benchmarks: recorded runs as ground truth, code-backed grades, and independent review.
Document / current
The expert’s own recorded run is the ground truth. A check that cannot recognize that run as correct does not ship.
Source
Working practitioners, identity-verified and paid at every stage, describe the work and then perform it on the real software. They review the task definition themselves. Nothing enters a corpus on a model’s word alone.
Grade
Wherever the software allows it, the grade is code: database state, ledgers, files, audit logs. Two rules are absolute. The verifier must reproduce the authoring expert’s recorded run. Ground truth is a set. When two independent runs diverge, the divergence is captured as tolerance.
Independence
No expert grades their own work. One authors, a second runs the task blind, a third adjudicates. Public cuts carry a leaderboard and a harness. Held-out sets never publish.
Request supply