Environment Bundle
Off-The-Shelf Data Catalog
Pre-built commercial datasets and RL environments spanning coding, enterprise agents, STEM, tool use, and reasoning.
Surge · 100% completeCompany-reported
A benchmark for evaluating AI systems that generate modern web applications.
App-Bench evaluates coding agents on building web applications. EnvIndex currently records only the fields supported by AfterQuery’s public benchmark page; task counts and technical access mechanisms remain unknown here.
Only evidenced capabilities are shown as Yes. Missing evidence remains Unknown.
Pre-built commercial datasets and RL environments spanning coding, enterprise agents, STEM, tool use, and reasoning.
A long-horizon coding eval in which agents build a Game Boy Advance emulator in Rust within 24 hours.
A long-horizon software-engineering benchmark built from public and proprietary repositories.