RL Environments and Agents
Custom RL environments and verifier design for training and evaluating agentic models.
Platforms that build environments, evals, and benchmarks across many domains at once: generalist environment foundries serving frontier labs with everything from coding to customer support. Scale is their moat: thousands of environments, expert networks in every field, and pipelines that turn any workflow into a training ground.
13 companies · 11 cataloged artifacts · as of 2026-08-16
| Name | Domains | Type / catalog | Location | Team | Funding | Raising |
|---|---|---|---|---|---|---|
| AfterQuery Expert human data and RL environments across code, finance, and computer use | Multi-Domain, Coding, Finance | Pure-play commercial 1 artifacts | San Francisco, New York, Seattle | 51-100 | $30.5M total (reported) | Unknown |
| Artificial Analysis Independent benchmarking of AI models across intelligence, speed, and price | Multi-Domain, Machine Learning | Pure-play commercial Catalog unknown | San Francisco | 11-25 | $2.6M (2024) | Unknown |
| Handshake Career network turned human-data and RL environments provider via Handshake AI | Multi-Domain, Data Labeling, RLHF | Data + environments 2 artifacts | San Francisco, New York, Bangalore, Berlin | 250+ | Series F, $200M (2022, ~$3.5B valuation) | Unknown |
| LMArena Crowdsourced model leaderboards from the Chatbot Arena team | Multi-Domain, Machine Learning | Pure-play commercial Catalog unknown | San Francisco, Berkeley | 26-50 | $100M seed (a16z, UC Investments, 2025); $150M at $1.7B valuation (Jan 2026) | Unknown |
| Mercor Expert marketplace powering evals and RL environments for frontier labs | Multi-Domain, Data Labeling, RLHF | Acquired / inactive 1 artifacts | San Francisco | 250+ | Series C, $350M at $10B valuation (Oct 2025) | Yes |
| Micro1 Vetted domain experts for AI training data and evals | Data Labeling, Multi-Domain, RLHF | Data + environments Catalog unknown | Los Angeles | 101-250 | Series A, $35M at $500M valuation (Sept 2025) | Unknown |
| Pareto Expert data workforce for RLHF, evals, and tool-use environments | Multi-Domain, Tool Use, Data Labeling, RLHF | Data + environments Catalog unknown | San Francisco | 26-50 | - | Unknown |
| Prime Intellect Open superintelligence stack: compute, RL environments hub, and sandboxes | Machine Learning, Environment Platforms, Multi-Domain | Environment platform Catalog unknown | San Francisco | 26-50 | Series A, $130M at $1B valuation (2026); $150M+ total | Unknown |
| Rise Data Labs RL environments and tasks across multiple domains, supported by a large expert network | Computer Use, Coding, Finance, Cybersecurity, Legal, Data Labeling, Enterprise, Multi-Domain, RLHF, Custom Environments | Data + environments 3 artifacts | New York, United States | 11-25 | - | Unknown |
| Scale Data-labeling incumbent extending into agent evals and RL environments | Multi-Domain, Coding, Data Labeling, RLHF | Data + environments 2 artifacts | San Francisco, New York, Washington DC, London | 250+ | $1.6B+ raised; Meta invested $14.3B at ~$29B valuation (June 2025) | Unknown |
| Snorkel Programmatic data platform expanding into expert evals and RL environments | Multi-Domain, Coding, Machine Learning, Data Labeling, RLHF | Data + environments Catalog unknown | San Francisco, Redwood City, New York | 101-250 | Series D, $100M at $1.3B valuation (2025) | Unknown |
| Surge Bootstrapped human-data leader with a dedicated RL environments org | Multi-Domain, Data Labeling, RLHF | Data + environments 11 artifacts | San Francisco, New York, Seattle | 101-250 | Bootstrapped; reported in talks to raise ~$1B at $25B+ valuation (2025) | Yes |
| Turing AGI infrastructure: coding data and RL environments at scale | Coding, Multi-Domain, Data Labeling, RLHF | Data + environments Catalog unknown | San Francisco, Palo Alto, Gurugram | 250+ | Series E, $111M at $2.2B valuation (2025) | Unknown |
Source-backed catalog records connected to multi-domain.
Custom RL environments and verifier design for training and evaluating agentic models.
Custom rubric and verifier design for scoring complex model and agent behavior.
Human preference and reward data for reinforcement learning from human feedback.
Expert demonstrations for bootstrapping model capabilities, including computer and browser use.
Human evaluation programs for model quality, usefulness, safety, and subjective output characteristics.
Expert-authored data and judgment across professional, STEM, and humanities domains.
Language and culturally grounded training data across more than 70 reported languages.
Training and evaluation data spanning text, images, audio, and video.
Pre-built commercial datasets and RL environments spanning coding, enterprise agents, STEM, tool use, and reasoning.
Private commercial long-horizon tasks spanning multiple knowledge-work domains.
A custom environment-development capability based on a buyer’s workflows, tools, data, and success criteria.