Off-The-Shelf Data Catalog
Pre-built commercial datasets and RL environments spanning coding, enterprise agents, STEM, tool use, and reasoning.
Environments that train and evaluate models on real software engineering: resolving issues in production repositories, writing and reviewing pull requests, debugging failing test suites, and shipping features end-to-end. Coding is the most mature RL environment category: benchmarks like SWE-bench proved that verifiable rewards from test suites scale, and frontier labs now buy coding environments by the thousand.
23 companies · 5 cataloged artifacts · as of 2026-08-16
| Name | Domains | Type / catalog | Location | Team | Funding | Raising |
|---|---|---|---|---|---|---|
| AfterQuery Expert human data and RL environments across code, finance, and computer use | Multi-Domain, Coding, Finance | Pure-play commercial 1 artifacts | San Francisco, New York, Seattle | 51-100 | $30.5M total (reported) | Unknown |
| Akhara Enterprise and code RL environments | Enterprise, Coding | Pure-play commercial Catalog unknown | San Francisco | 1-10 | - | Unknown |
| BenchFlow Open-source benchmark hub and eval infrastructure for agents | Enterprise, Browser, Coding | Pure-play commercial Catalog unknown | San Francisco | 1-10 | ~$1M (reported) | Unknown |
| Bespoke Labs Data curation and RL environment recipes from ex-Google DeepMind researchers | Coding, Machine Learning | Pure-play commercial Catalog unknown | Mountain View, Menlo Park, Bangalore, San Francisco | 11-25 | ~$40M (reported) | Unknown |
| Datacurve Frontier coding data and repository RL environments via the Shipd bounty platform | Coding, RLHF | Data + environments Catalog unknown | San Francisco | 26-50 | Series A, $15M led by Chemistry (Oct 2025); $17.7M total | Unknown |
| Deeptune Code and computer-use environments; acquired by Mercor | Coding, Computer Use | Acquired / inactive Catalog unknown | New York | 26-50 | Series A, $43M; acquired by Mercor (2026) | Unknown |
| Exabite Code RL environments with realistic software execution | Coding | Pure-play commercial Catalog unknown | Remote | 1-10 | - | Unknown |
| HUD Evals and RL environments platform for computer-use agents | Computer Use, Coding, Long Horizon, Environment Platforms, Custom Environments | Environment platform Catalog unknown | San Francisco, Singapore | 11-25 | $15M raised (YC W25, Exceptional Capital) | Unknown |
| Huzzle Labs Long-horizon code, tool-use, and enterprise workflow environments | Long Horizon, Coding, Enterprise | Pure-play commercial Catalog unknown | London, Berlin, San Francisco | 26-50 | ~$6M (reported) | Unknown |
| Idler Code environments with realistic execution constraints | Coding | Pure-play commercial Catalog unknown | San Francisco | 11-25 | - | Unknown |
| Mechanize RL environments to automate software engineering, founded by ex-Epoch AI researchers | Coding | Pure-play commercial 1 artifacts | San Francisco | 51-100 | ~$9.1M (reported) | Unknown |
| Metaphi Code and enterprise RL environments | Coding, Enterprise | Pure-play commercial Catalog unknown | San Francisco, New York | 1-10 | - | Unknown |
| pre.dev Software-planning platform offering coding and long-horizon RL environments | Coding, Long Horizon | Pure-play commercial Catalog unknown | Delaware | 1-10 | - | Unknown |
| Preference Model Stealth startup working on preference and reward modeling | Machine Learning, Coding | Pure-play commercial Catalog unknown | San Francisco, Toronto, Seattle | 11-25 | - | Unknown |
| Proximal Long-horizon coding RL environments built from real codebases | Coding, Long Horizon | Pure-play commercial Catalog unknown | San Francisco, Bangalore | 26-50 | - | Unknown |
| Quesma Security-domain RL environments and binary analysis evals | Cybersecurity, Coding | Pure-play commercial Catalog unknown | Warsaw | 11-25 | - | Unknown |
| ReasonCore Science and code reasoning environments and benchmarks | Science, Coding | Pure-play commercial Catalog unknown | San Francisco | 11-25 | - | Unknown |
| Refresh Simulation engines with verifiable rewards for coding and computer use | Coding, Computer Use, Simulation | Pure-play commercial Catalog unknown | San Francisco | 1-10 | - | Unknown |
| Rise Data Labs RL environments and tasks across multiple domains, supported by a large expert network | Computer Use, Coding, Finance, Cybersecurity, Legal, Data Labeling, Enterprise, Multi-Domain, RLHF, Custom Environments | Data + environments 3 artifacts | New York, United States | 11-25 | - | Unknown |
| Scale Data-labeling incumbent extending into agent evals and RL environments | Multi-Domain, Coding, Data Labeling, RLHF | Data + environments 2 artifacts | San Francisco, New York, Washington DC, London | 250+ | $1.6B+ raised; Meta invested $14.3B at ~$29B valuation (June 2025) | Unknown |
| Snorkel Programmatic data platform expanding into expert evals and RL environments | Multi-Domain, Coding, Machine Learning, Data Labeling, RLHF | Data + environments Catalog unknown | San Francisco, Redwood City, New York | 101-250 | Series D, $100M at $1.3B valuation (2025) | Unknown |
| Turing AGI infrastructure: coding data and RL environments at scale | Coding, Multi-Domain, Data Labeling, RLHF | Data + environments Catalog unknown | San Francisco, Palo Alto, Gurugram | 250+ | Series E, $111M at $2.2B valuation (2025) | Unknown |
| Vetto AI Code and computer-use environments from ex-DeepMind/Instagram founders | Coding, Computer Use | Pure-play commercial Catalog unknown | San Francisco, São Paulo, London | 1-10 | - | Unknown |
Source-backed catalog records connected to coding.
Pre-built commercial datasets and RL environments spanning coding, enterprise agents, STEM, tool use, and reasoning.
A long-horizon coding eval in which agents build a Game Boy Advance emulator in Rust within 24 hours.
A long-horizon software-engineering benchmark built from public and proprietary repositories.
The proprietary-repository subset of SWE-Bench Pro, documented publicly but not distributed as an open dataset.
A benchmark for evaluating AI systems that generate modern web applications.