RL Environments and Agents
Custom RL environments and verifier design for training and evaluating agentic models.
Canonical market-intelligence pages for the artifacts vendors and research teams actually offer. Public availability, commercial access, and verification are tracked separately.
21 of 21 artifacts
Custom RL environments and verifier design for training and evaluating agentic models.
Custom rubric and verifier design for scoring complex model and agent behavior.
Human preference and reward data for reinforcement learning from human feedback.
Expert demonstrations for bootstrapping model capabilities, including computer and browser use.
Human evaluation programs for model quality, usefulness, safety, and subjective output characteristics.
Expert-authored data and judgment across professional, STEM, and humanities domains.
Language and culturally grounded training data across more than 70 reported languages.
Training and evaluation data spanning text, images, audio, and video.
Pre-built commercial datasets and RL environments spanning coding, enterprise agents, STEM, tool use, and reasoning.
A simulated enterprise environment where agents complete customer-support and operational tasks inside a fictional PC retailer.
A long-context enterprise-agent benchmark built from unique RL environments with internal tools and external MCP servers.
Private commercial long-horizon tasks spanning multiple knowledge-work domains.
A private commercial catalog of finance tasks for RL training and evaluation.
A custom environment-development capability based on a buyer’s workflows, tools, data, and success criteria.
A long-horizon coding eval in which agents build a Game Boy Advance emulator in Rust within 24 hours.
A long-horizon software-engineering benchmark built from public and proprietary repositories.
The proprietary-repository subset of SWE-Bench Pro, documented publicly but not distributed as an open dataset.
An open-source reactive agent-as-judge that inspects files and environment state while grading agent work.
A banking-workflow benchmark used to evaluate agents and verifier behavior on artifact-heavy tasks.
A benchmark for evaluating AI systems that generate modern web applications.
An expert-authored benchmark of economically valuable tasks in finance, consulting, law, and medicine.