RLHF
RLHF vendors: companies that run human-in-the-loop pipelines for reinforcement learning from human feedback, from preference data and human evaluation to expert-graded rewards. Many offer RLHF as a managed service alongside evals, benchmarks, and RL environments, recruiting domain experts to rate, rank, and correct model outputs at scale.
13 startups
| Name | Domains | Location | Team | Funding | Raising |
|---|---|---|---|---|---|
| Besimple AI Human-in-the-loop annotation and evals, with a voice data focus | Data Labeling, Voice, RLHF | San Francisco | 1-10 | ~$3.5M total | - |
| Datacurve Frontier coding data and repository RL environments via the Shipd bounty platform | Coding, RLHF | San Francisco | 26-50 | Series A, $15M led by Chemistry (Oct 2025); $17.7M total | - |
| Handshake Career network turned human-data and RL environments provider via Handshake AI | Multi-Domain, Data Labeling, RLHF | San Francisco, New York, Bangalore, Berlin | 250+ | Series F, $200M (2022, ~$3.5B valuation) | - |
| Mercor Expert marketplace powering evals and RL environments for frontier labs | Multi-Domain, Data Labeling, RLHF | San Francisco | 250+ | Series C, $350M at $10B valuation (Oct 2025) | Yes |
| Micro1 Vetted domain experts for AI training data and evals | Data Labeling, Multi-Domain, RLHF | Los Angeles | 101-250 | Series A, $35M at $500M valuation (Sept 2025) | - |
| Pareto Expert data workforce for RLHF, evals, and tool-use environments | Multi-Domain, Tool Use, Data Labeling, RLHF | San Francisco | 26-50 | - | - |
| Rise Data Labs US-based expert human data and custom RL environments for enterprise AI | Computer Use, Coding, Finance, Cybersecurity, Legal, Data Labeling, Enterprise, Multi-Domain, RLHF, Custom Environments | United States | - | - | Yes |
| Scale Data-labeling incumbent extending into agent evals and RL environments | Multi-Domain, Coding, Data Labeling, RLHF | San Francisco, New York, Washington DC, London | 250+ | $1.6B+ raised; Meta invested $14.3B at ~$29B valuation (June 2025) | - |
| Snorkel Programmatic data platform expanding into expert evals and RL environments | Multi-Domain, Coding, Machine Learning, Data Labeling, RLHF | San Francisco, Redwood City, New York | 101-250 | Series D, $100M at $1.3B valuation (2025) | - |
| Surge Bootstrapped human-data leader with a dedicated RL environments org | Multi-Domain, Data Labeling, RLHF | San Francisco, New York, Seattle | 101-250 | Bootstrapped; reported in talks to raise ~$1B at $25B+ valuation (2025) | Yes |
| SynthLabs Post-training research: synthetic data and scalable RL alignment | Machine Learning, Alignment, RLHF | San Francisco | 1-10 | Seed (M12 and First Spark Ventures, 2024) | - |
| The LLM Data Company Evals and reward data tooling for LLM training | Machine Learning, Data Labeling, RLHF | - | 1-10 | - | - |
| Turing AGI infrastructure: coding data and RL environments at scale | Coding, Multi-Domain, Data Labeling, RLHF | San Francisco, Palo Alto, Gurugram | 250+ | Series E, $111M at $2.2B valuation (2025) | - |