rlsupply · research supply for reinforcement learning environments / verified

Method

Methodology

How RLSupply turns expert work into environments and benchmarks: recorded runs as ground truth, code-backed grades, and independent review.

Document / current

The expert’s own recorded run is the ground truth. A check that cannot recognize that run as correct does not ship.

Source

Working practitioners, identity-verified and paid at every stage, describe the work and then perform it on the real software. They review the task definition themselves. Nothing enters a corpus on a model’s word alone.

Grade

Wherever the software allows it, the grade is code: database state, ledgers, files, audit logs. Two rules are absolute. The verifier must reproduce the authoring expert’s recorded run. Ground truth is a set. When two independent runs diverge, the divergence is captured as tolerance.

Independence

No expert grades their own work. One authors, a second runs the task blind, a third adjudicates. Public cuts carry a leaderboard and a harness. Held-out sets never publish.

Request supply

Tell us where your model breaks.
We build the environment.