Benchmark
SWE-Bench Pro - Private Dataset
The proprietary-repository subset of SWE-Bench Pro, documented publicly but not distributed as an open dataset.
Scale · 100% completeCompany-reported
A long-horizon software-engineering benchmark built from public and proprietary repositories.
SWE-Bench Pro evaluates agents on realistic software-engineering problems. The published paper reports 1,865 problems from 41 actively maintained repositories; the public dataset and leaderboard are separately accessible.
Only evidenced capabilities are shown as Yes. Missing evidence remains Unknown.
The proprietary-repository subset of SWE-Bench Pro, documented publicly but not distributed as an open dataset.
Custom RL environments and verifier design for training and evaluating agentic models.
Pre-built commercial datasets and RL environments spanning coding, enterprise agents, STEM, tool use, and reasoning.