Benchmark
CoreCraft
A simulated enterprise environment where agents complete customer-support and operational tasks inside a fictional PC retailer.
Surge · 100% completeCompany-reported
A long-context enterprise-agent benchmark built from unique RL environments with internal tools and external MCP servers.
HANDBOOK.md evaluates whether agents can follow a long company handbook while completing day-to-day professional tasks. Surge reports that each task is a unique RL environment across five enterprise domains.
Only evidenced capabilities are shown as Yes. Missing evidence remains Unknown.
A simulated enterprise environment where agents complete customer-support and operational tasks inside a fictional PC retailer.
Custom RL environments and verifier design for training and evaluating agentic models.
Pre-built commercial datasets and RL environments spanning coding, enterprise agents, STEM, tool use, and reasoning.