Benchmark·Public sample

App-Bench

A benchmark for evaluating AI systems that generate modern web applications.

Company-reportedConfidence: medium·Last verified: 2026-08-16·How verification works
Access
Public / Open
Profile completeness 100%

What it is

App-Bench evaluates coding agents on building web applications. EnvIndex currently records only the fields supported by AfterQuery’s public benchmark page; task counts and technical access mechanisms remain unknown here.

Capabilities

Only evidenced capabilities are shown as Yes. Missing evidence remains Unknown.

StatefulUnknown
Long horizonUnknown
Multi-turnUnknown
Multi-agentUnknown
AdversarialUnknown
MultimodalUnknown
Tool useUnknown
Browser useUnknown
Computer useUnknown
Code executionYes
File manipulationUnknown
API interactionUnknown
MCP supportUnknown
SandboxedUnknown
Deterministic resetUnknown
Stochastic behaviorUnknown
Human in the loopUnknown
Expert authoredUnknown
SyntheticUnknown
Production derivedUnknown
Real application replicaUnknown
Partial observabilityUnknown
Persistent stateUnknown
Cross-applicationUnknown

Sources & provenance

  1. AfterQuery · official · accessed 2026-08-16 · company-reported

Related artifacts

Benchmark

GBA Eval

Public sample

A long-horizon coding eval in which agents build a Game Boy Advance emulator in Rust within 24 hours.

Mechanize · 100% completeCompany-reported
Benchmark

SWE-Bench Pro

Public sample

A long-horizon software-engineering benchmark built from public and proprietary repositories.

Scale · 100% completeConfirmed