RL Environments and Agents
Custom RL environments and verifier design for training and evaluating agentic models.
Replicas of the software that runs companies (CRMs, ERPs, ticketing systems, spreadsheets, email) wired into environments where agents complete real business workflows. The bet: the biggest economic value of agents is in enterprise back-office work, so labs need training grounds that look exactly like it.
10 companies · 11 cataloged artifacts · as of 2026-08-16
| Name | Domains | Type / catalog | Location | Team | Funding | Raising |
|---|---|---|---|---|---|---|
| Akhara Enterprise and code RL environments | Enterprise, Coding | Pure-play commercial Catalog unknown | San Francisco | 1-10 | - | Unknown |
| Applied Compute Ex-OpenAI trio applying RL to build specialist enterprise models | Enterprise, Machine Learning, Custom Environments | Pure-play commercial Catalog unknown | San Francisco | 11-25 | $80M total at $700M valuation (Oct 2025) | Yes |
| BenchFlow Open-source benchmark hub and eval infrastructure for agents | Enterprise, Browser, Coding | Pure-play commercial Catalog unknown | San Francisco | 1-10 | ~$1M (reported) | Unknown |
| Collinear Enterprise simulation, judges, and long-horizon trajectory generation | Enterprise, Long Horizon, Machine Learning, Simulation | Pure-play commercial Catalog unknown | Mountain View, Sunnyvale | 11-25 | - | Unknown |
| Halluminate Sandboxed RL environments for finance and enterprise workflows | Finance, Enterprise, Browser | Pure-play commercial Catalog unknown | San Francisco | 26-50 | - | Unknown |
| Huzzle Labs Long-horizon code, tool-use, and enterprise workflow environments | Long Horizon, Coding, Enterprise | Pure-play commercial Catalog unknown | London, Berlin, San Francisco | 26-50 | ~$6M (reported) | Unknown |
| Metaphi Code and enterprise RL environments | Coding, Enterprise | Pure-play commercial Catalog unknown | San Francisco, New York | 1-10 | - | Unknown |
| Plato High-fidelity replicas of websites and software for agent training | Browser, Enterprise, Simulation | Pure-play commercial Catalog unknown | San Francisco | 1-10 | - | Unknown |
| Rise Data Labs RL environments and tasks across multiple domains, supported by a large expert network | Computer Use, Coding, Finance, Cybersecurity, Legal, Data Labeling, Enterprise, Multi-Domain, RLHF, Custom Environments | Data + environments 3 artifacts | New York, United States | 11-25 | - | Unknown |
| Veris AI High-fidelity simulated environments to train enterprise AI agents | Enterprise, Custom Environments, Simulation | Pure-play commercial Catalog unknown | New York | 1-10 | Seed, $8.5M (Decibel and Acrew, June 2025) | Unknown |
Source-backed catalog records connected to enterprise.
Custom RL environments and verifier design for training and evaluating agentic models.
Pre-built commercial datasets and RL environments spanning coding, enterprise agents, STEM, tool use, and reasoning.
A simulated enterprise environment where agents complete customer-support and operational tasks inside a fictional PC retailer.
A long-context enterprise-agent benchmark built from unique RL environments with internal tools and external MCP servers.
Private commercial long-horizon tasks spanning multiple knowledge-work domains.
A private commercial catalog of finance tasks for RL training and evaluation.
A custom environment-development capability based on a buyer’s workflows, tools, data, and success criteria.
The proprietary-repository subset of SWE-Bench Pro, documented publicly but not distributed as an open dataset.
An open-source reactive agent-as-judge that inspects files and environment state while grading agent work.
A banking-workflow benchmark used to evaluate agents and verifier behavior on artifact-heavy tasks.
An expert-authored benchmark of economically valuable tasks in finance, consulting, law, and medicine.