Dataset
RLHF
Human preference and reward data for reinforcement learning from human feedback.
Surge · 100% completeCompany-reported
RLHF vendors: companies that run human-in-the-loop pipelines for reinforcement learning from human feedback, from preference data and human evaluation to expert-graded rewards. Many offer RLHF as a managed service alongside evals, benchmarks, and RL environments, recruiting domain experts to rate, rank, and correct model outputs at scale.
11 companies · 1 cataloged artifact · as of 2026-08-16
| Name | Domains | Type / catalog | Location | Team | Funding | Raising |
|---|---|---|---|---|---|---|
| Datacurve Frontier coding data and repository RL environments via the Shipd bounty platform | Coding, RLHF | Data + environments Catalog unknown | San Francisco | 26-50 | Series A, $15M led by Chemistry (Oct 2025); $17.7M total | Unknown |
| Handshake Career network turned human-data and RL environments provider via Handshake AI | Multi-Domain, Data Labeling, RLHF | Data + environments 2 artifacts | San Francisco, New York, Bangalore, Berlin | 250+ | Series F, $200M (2022, ~$3.5B valuation) | Unknown |
| Mercor Expert marketplace powering evals and RL environments for frontier labs | Multi-Domain, Data Labeling, RLHF | Acquired / inactive 1 artifacts | San Francisco | 250+ | Series C, $350M at $10B valuation (Oct 2025) | Yes |
| Micro1 Vetted domain experts for AI training data and evals | Data Labeling, Multi-Domain, RLHF | Data + environments Catalog unknown | Los Angeles | 101-250 | Series A, $35M at $500M valuation (Sept 2025) | Unknown |
| Pareto Expert data workforce for RLHF, evals, and tool-use environments | Multi-Domain, Tool Use, Data Labeling, RLHF | Data + environments Catalog unknown | San Francisco | 26-50 | - | Unknown |
| Rise Data Labs RL environments and tasks across multiple domains, supported by a large expert network | Computer Use, Coding, Finance, Cybersecurity, Legal, Data Labeling, Enterprise, Multi-Domain, RLHF, Custom Environments | Data + environments 3 artifacts | New York, United States | 11-25 | - | Unknown |
| Scale Data-labeling incumbent extending into agent evals and RL environments | Multi-Domain, Coding, Data Labeling, RLHF | Data + environments 2 artifacts | San Francisco, New York, Washington DC, London | 250+ | $1.6B+ raised; Meta invested $14.3B at ~$29B valuation (June 2025) | Unknown |
| Snorkel Programmatic data platform expanding into expert evals and RL environments | Multi-Domain, Coding, Machine Learning, Data Labeling, RLHF | Data + environments Catalog unknown | San Francisco, Redwood City, New York | 101-250 | Series D, $100M at $1.3B valuation (2025) | Unknown |
| Surge Bootstrapped human-data leader with a dedicated RL environments org | Multi-Domain, Data Labeling, RLHF | Data + environments 11 artifacts | San Francisco, New York, Seattle | 101-250 | Bootstrapped; reported in talks to raise ~$1B at $25B+ valuation (2025) | Yes |
| SynthLabs Post-training research: synthetic data and scalable RL alignment | Machine Learning, Alignment, RLHF | Data + environments Catalog unknown | San Francisco | 1-10 | Seed (M12 and First Spark Ventures, 2024) | Unknown |
| Turing AGI infrastructure: coding data and RL environments at scale | Coding, Multi-Domain, Data Labeling, RLHF | Data + environments Catalog unknown | San Francisco, Palo Alto, Gurugram | 250+ | Series E, $111M at $2.2B valuation (2025) | Unknown |
Source-backed catalog records connected to rlhf.
Human preference and reward data for reinforcement learning from human feedback.