RL Environments and Agents
Custom RL environments and verifier design for training and evaluating agentic models.
Bootstrapped human-data leader with a dedicated RL environments org
Surge AI is the largest human data / RLHF vendor by revenue ($1.2B+ in 2024) serving OpenAI, Google, Anthropic, and Meta, and has created an internal organization dedicated to RL environments. It builds expert-graded evals and environment products such as EnterpriseBench.
~$9.9M revenue per employee; CoreCraft benchmark
| Website | surgehq.ai |
|---|---|
| Domains | Multi-Domain, Data Labeling, RLHF |
| Location | San Francisco, New York, Seattle |
| Team size | 101-250 |
| Founded | 2020 |
| Funding | Bootstrapped; reported in talks to raise ~$1B at $25B+ valuation (2025) |
| Raising | Reported yes · re-verification required |
| Founders | Edwin Chen |
Focus areas and technical capabilities are shown separately. Missing technical evidence remains Unknown.
Products and services supported by official company materials. Reviewed 2026-08-16.
Custom RL environments and verifier design for training and evaluating agentic models.
Custom rubric and verifier design for scoring complex model and agent behavior.
Human preference and reward data for reinforcement learning from human feedback.
Expert demonstrations for bootstrapping model capabilities, including computer and browser use.
Human evaluation programs for model quality, usefulness, safety, and subjective output characteristics.
Expert-authored data and judgment across professional, STEM, and humanities domains.
Language and culturally grounded training data across more than 70 reported languages.
Training and evaluation data spanning text, images, audio, and video.
Pre-built commercial datasets and RL environments spanning coding, enterprise agents, STEM, tool use, and reasoning.
A simulated enterprise environment where agents complete customer-support and operational tasks inside a fictional PC retailer.
A long-context enterprise-agent benchmark built from unique RL environments with internal tools and external MCP servers.
Verified public artifacts and documented private commercial inventory.
Custom RL environments and verifier design for training and evaluating agentic models.
Custom rubric and verifier design for scoring complex model and agent behavior.
Human preference and reward data for reinforcement learning from human feedback.
Expert demonstrations for bootstrapping model capabilities, including computer and browser use.
Human evaluation programs for model quality, usefulness, safety, and subjective output characteristics.
Expert-authored data and judgment across professional, STEM, and humanities domains.
Language and culturally grounded training data across more than 70 reported languages.
Training and evaluation data spanning text, images, audio, and video.
Pre-built commercial datasets and RL environments spanning coding, enterprise agents, STEM, tool use, and reasoning.
A simulated enterprise environment where agents complete customer-support and operational tasks inside a fictional PC retailer.
A long-context enterprise-agent benchmark built from unique RL environments with internal tools and external MCP servers.
Company representatives and researchers can propose sourced changes. Submissions do not directly overwrite editorial data.
Also listed under Multi-Domain, Data Labeling, RLHF.
| Name | Domains | Type / catalog | Location | Team | Funding | Raising |
|---|---|---|---|---|---|---|
| Rise Data Labs RL environments and tasks across multiple domains, supported by a large expert network | Computer Use, Coding, Finance, Cybersecurity, Legal, Data Labeling, Enterprise, Multi-Domain, RLHF, Custom Environments | Data + environments 3 artifacts | New York, United States | 11-25 | - | Unknown |
| AfterQuery Expert human data and RL environments across code, finance, and computer use | Multi-Domain, Coding, Finance | Pure-play commercial 1 artifacts | San Francisco, New York, Seattle | 51-100 | $30.5M total (reported) | Unknown |
| Anthromind Medical and long-horizon environments and expert data | Medical, Long Horizon, Data Labeling | Data + environments Catalog unknown | San Francisco | 1-10 | - | Unknown |
| Artificial Analysis Independent benchmarking of AI models across intelligence, speed, and price | Multi-Domain, Machine Learning | Pure-play commercial Catalog unknown | San Francisco | 11-25 | $2.6M (2024) | Unknown |
| Datacurve Frontier coding data and repository RL environments via the Shipd bounty platform | Coding, RLHF | Data + environments Catalog unknown | San Francisco | 26-50 | Series A, $15M led by Chemistry (Oct 2025); $17.7M total | Unknown |
| Handshake Career network turned human-data and RL environments provider via Handshake AI | Multi-Domain, Data Labeling, RLHF | Data + environments 2 artifacts | San Francisco, New York, Bangalore, Berlin | 250+ | Series F, $200M (2022, ~$3.5B valuation) | Unknown |