Benchmark
CoreCraft
A simulated enterprise environment where agents complete customer-support and operational tasks inside a fictional PC retailer.
Surge · 100% completeCompany-reported
Custom RL environments and verifier design for training and evaluating agentic models.
Surge describes a capability for creating complex RL environments that challenge agentic models, together with verifiers that reward their behavior. This capability record does not imply a fixed public task count or a single packaged environment.
Only evidenced capabilities are shown as Yes. Missing evidence remains Unknown.
A simulated enterprise environment where agents complete customer-support and operational tasks inside a fictional PC retailer.
A long-context enterprise-agent benchmark built from unique RL environments with internal tools and external MCP servers.
Custom rubric and verifier design for scoring complex model and agent behavior.