Benchmark
SWE-Bench Pro
A long-horizon software-engineering benchmark built from public and proprietary repositories.
Scale · 100% completeConfirmed
The proprietary-repository subset of SWE-Bench Pro, documented publicly but not distributed as an open dataset.
This private subset evaluates coding agents against commercial-grade proprietary repositories. Its existence and leaderboard are public; the underlying repository contents are not presented as publicly downloadable by EnvIndex.
Only evidenced capabilities are shown as Yes. Missing evidence remains Unknown.
A long-horizon software-engineering benchmark built from public and proprietary repositories.
Custom RL environments and verifier design for training and evaluating agentic models.
Pre-built commercial datasets and RL environments spanning coding, enterprise agents, STEM, tool use, and reasoning.