RL Environments and Agents
Custom RL environments and verifier design for training and evaluating agentic models.
Environments testing whether agents can wield tools: chaining APIs, recovering from tool errors, choosing the right function among hundreds, and composing multi-tool workflows. As MCP-style ecosystems grow, tool-use environments are how labs keep models reliable across thousands of integrations.
2 companies · 6 cataloged artifacts · as of 2026-08-16
| Name | Domains | Type / catalog | Location | Team | Funding | Raising |
|---|---|---|---|---|---|---|
| Chakra Labs Dojo: a hub of computer-use and tool-use environments | Computer Use, Tool Use | Pure-play commercial Catalog unknown | Brooklyn | 11-25 | ~$10.1M (reported) | Unknown |
| Pareto Expert data workforce for RLHF, evals, and tool-use environments | Multi-Domain, Tool Use, Data Labeling, RLHF | Data + environments Catalog unknown | San Francisco | 26-50 | - | Unknown |
Source-backed catalog records connected to tool use.
Custom RL environments and verifier design for training and evaluating agentic models.
Pre-built commercial datasets and RL environments spanning coding, enterprise agents, STEM, tool use, and reasoning.
A simulated enterprise environment where agents complete customer-support and operational tasks inside a fictional PC retailer.
A long-context enterprise-agent benchmark built from unique RL environments with internal tools and external MCP servers.
An open-source reactive agent-as-judge that inspects files and environment state while grading agent work.
A banking-workflow benchmark used to evaluate agents and verifier behavior on artifact-heavy tasks.