RL Environments and Agents
Custom RL environments and verifier design for training and evaluating agentic models.
Environments deliberately designed to take hours or days of agent effort: sprawling codebases, multi-stage projects, tasks with delayed and sparse rewards. Long-horizon coherence is the axis on which current frontier models improve fastest, and these environments are how that improvement gets measured and trained.
14 companies · 8 cataloged artifacts · as of 2026-08-16
| Name | Domains | Type / catalog | Location | Team | Funding | Raising |
|---|---|---|---|---|---|---|
| Andon Labs Long-horizon autonomy benchmarks like Vending-Bench | Long Horizon, Alignment | Pure-play commercial Catalog unknown | San Francisco | 1-10 | Seed (Y Combinator) | Unknown |
| Andromede Programmatic generation of long-horizon RL environments | Long Horizon | Pure-play commercial Catalog unknown | Lausanne | 1-10 | - | Unknown |
| Anthromind Medical and long-horizon environments and expert data | Medical, Long Horizon, Data Labeling | Data + environments Catalog unknown | San Francisco | 1-10 | - | Unknown |
| ARIMLABS Security and long-horizon environments for agentic AI | Cybersecurity, Long Horizon | Pure-play commercial Catalog unknown | Warsaw | 11-25 | - | Unknown |
| Collinear Enterprise simulation, judges, and long-horizon trajectory generation | Enterprise, Long Horizon, Machine Learning, Simulation | Pure-play commercial Catalog unknown | Mountain View, Sunnyvale | 11-25 | - | Unknown |
| Diffuse Labs ML and long-horizon RL environments | Machine Learning, Long Horizon | Pure-play commercial Catalog unknown | Palo Alto, San Francisco | 1-10 | - | Unknown |
| EdotEnv Long-horizon planning environments for frontier models | Long Horizon, Machine Learning | Pure-play commercial Catalog unknown | San Francisco | 1-10 | - | Unknown |
| General Reasoning Open reasoning data and reward models from the ex-Meta AI reasoning lead | Finance, Long Horizon, Machine Learning | Pure-play commercial Catalog unknown | London, San Francisco | 1-10 | ~$10.9M (reported) | Unknown |
| Good Start Labs Game-based RL environments and benchmarks | Games, Long Horizon | Pure-play commercial Catalog unknown | Brooklyn, New York, Toronto | 1-10 | ~$3.6M (reported) | Unknown |
| HUD Evals and RL environments platform for computer-use agents | Computer Use, Coding, Long Horizon, Environment Platforms, Custom Environments | Environment platform Catalog unknown | San Francisco, Singapore | 11-25 | $15M raised (YC W25, Exceptional Capital) | Unknown |
| Huzzle Labs Long-horizon code, tool-use, and enterprise workflow environments | Long Horizon, Coding, Enterprise | Pure-play commercial Catalog unknown | London, Berlin, San Francisco | 26-50 | ~$6M (reported) | Unknown |
| pre.dev Software-planning platform offering coding and long-horizon RL environments | Coding, Long Horizon | Pure-play commercial Catalog unknown | Delaware | 1-10 | - | Unknown |
| Proximal Long-horizon coding RL environments built from real codebases | Coding, Long Horizon | Pure-play commercial Catalog unknown | San Francisco, Bangalore | 26-50 | - | Unknown |
| Tacit Labs Life-science and long-horizon environments for AI models | Science, Long Horizon | Pure-play commercial Catalog unknown | San Francisco | 1-10 | - | Unknown |
Source-backed catalog records connected to long horizon.
Custom RL environments and verifier design for training and evaluating agentic models.
Pre-built commercial datasets and RL environments spanning coding, enterprise agents, STEM, tool use, and reasoning.
A simulated enterprise environment where agents complete customer-support and operational tasks inside a fictional PC retailer.
A long-context enterprise-agent benchmark built from unique RL environments with internal tools and external MCP servers.
Private commercial long-horizon tasks spanning multiple knowledge-work domains.
A long-horizon coding eval in which agents build a Game Boy Advance emulator in Rust within 24 hours.
A long-horizon software-engineering benchmark built from public and proprietary repositories.
The proprietary-repository subset of SWE-Bench Pro, documented publicly but not distributed as an open dataset.