Agentic Data and Evaluations
Data and evaluation programs for agentic AI systems.
Data-labeling incumbent extending into agent evals and RL environments
Scale AI is the original data-labeling powerhouse for AI labs and enterprises, now building RL environments and agent evaluation products under its agents and RL environments group. After Meta's 2025 investment and the departure of CEO Alexandr Wang, it lost some lab customers but continues to push into environments.
Publishes SWE-bench Pro benchmark
| Website | scale.com |
|---|---|
| Domains | Multi-Domain, Coding, Data Labeling, RLHF |
| Location | San Francisco, New York, Washington DC, London |
| Team size | 250+ |
| Founded | 2016 |
| Funding | $1.6B+ raised; Meta invested $14.3B at ~$29B valuation (June 2025) |
| Founders | Alexandr Wang @alexandr_wang, Lucy Guo |
Focus areas and technical capabilities are shown separately. Missing technical evidence remains Unknown.
Products and services supported by official company materials. Reviewed 2026-08-16.
Data and evaluation programs for agentic AI systems.
A long-horizon software-engineering benchmark built from public and proprietary repositories.
The proprietary-repository subset of SWE-Bench Pro, documented publicly but not distributed as an open dataset.
Verified public artifacts and documented private commercial inventory.
A long-horizon software-engineering benchmark built from public and proprietary repositories.
The proprietary-repository subset of SWE-Bench Pro, documented publicly but not distributed as an open dataset.
Company representatives and researchers can propose sourced changes. Submissions do not directly overwrite editorial data.
Also listed under Multi-Domain, Coding, Data Labeling, RLHF.
| Name | Domains | Type / catalog | Location | Team | Funding | Raising |
|---|---|---|---|---|---|---|
| Rise Data Labs RL environments and tasks across multiple domains, supported by a large expert network | Computer Use, Coding, Finance, Cybersecurity, Legal, Data Labeling, Enterprise, Multi-Domain, RLHF, Custom Environments | Data + environments 3 artifacts | New York, United States | 11-25 | - | Unknown |
| AfterQuery Expert human data and RL environments across code, finance, and computer use | Multi-Domain, Coding, Finance | Pure-play commercial 1 artifacts | San Francisco, New York, Seattle | 51-100 | $30.5M total (reported) | Unknown |
| Akhara Enterprise and code RL environments | Enterprise, Coding | Pure-play commercial Catalog unknown | San Francisco | 1-10 | - | Unknown |
| Anthromind Medical and long-horizon environments and expert data | Medical, Long Horizon, Data Labeling | Data + environments Catalog unknown | San Francisco | 1-10 | - | Unknown |
| Artificial Analysis Independent benchmarking of AI models across intelligence, speed, and price | Multi-Domain, Machine Learning | Pure-play commercial Catalog unknown | San Francisco | 11-25 | $2.6M (2024) | Unknown |
| BenchFlow Open-source benchmark hub and eval infrastructure for agents | Enterprise, Browser, Coding | Pure-play commercial Catalog unknown | San Francisco | 1-10 | ~$1M (reported) | Unknown |