GBA Eval
A long-horizon coding eval in which agents build a Game Boy Advance emulator in Rust within 24 hours.
RL environments to automate software engineering, founded by ex-Epoch AI researchers
Mechanize builds RL environments and 'boot camps' aimed at automating white-collar work, starting with software engineering agents. Founded by former Epoch AI leaders, it works with Anthropic on RL environments and is backed by investors including Nat Friedman, Daniel Gross, and Patrick Collison.
Offered $500K salaries to engineers building environments; publishes GBA Eval
| Website | mechanize.work |
|---|---|
| Domains | Coding |
| Location | San Francisco |
| Team size | 51-100 |
| Founded | 2025 |
| Funding | ~$9.1M (reported) |
| Founders | Tamay Besiroglu @tamaybes, Ege Erdil @EgeErdil2, Matthew Barnett @MatthewJBar |
Focus areas and technical capabilities are shown separately. Missing technical evidence remains Unknown.
Products and services supported by official company materials. Reviewed 2026-08-16.
A long-horizon coding eval in which agents build a Game Boy Advance emulator in Rust within 24 hours.
Verified public artifacts and documented private commercial inventory.
A long-horizon coding eval in which agents build a Game Boy Advance emulator in Rust within 24 hours.
Company representatives and researchers can propose sourced changes. Submissions do not directly overwrite editorial data.
Also listed under Coding.
| Name | Domains | Type / catalog | Location | Team | Funding | Raising |
|---|---|---|---|---|---|---|
| Rise Data Labs RL environments and tasks across multiple domains, supported by a large expert network | Computer Use, Coding, Finance, Cybersecurity, Legal, Data Labeling, Enterprise, Multi-Domain, RLHF, Custom Environments | Data + environments 3 artifacts | New York, United States | 11-25 | - | Unknown |
| AfterQuery Expert human data and RL environments across code, finance, and computer use | Multi-Domain, Coding, Finance | Pure-play commercial 1 artifacts | San Francisco, New York, Seattle | 51-100 | $30.5M total (reported) | Unknown |
| Akhara Enterprise and code RL environments | Enterprise, Coding | Pure-play commercial Catalog unknown | San Francisco | 1-10 | - | Unknown |
| BenchFlow Open-source benchmark hub and eval infrastructure for agents | Enterprise, Browser, Coding | Pure-play commercial Catalog unknown | San Francisco | 1-10 | ~$1M (reported) | Unknown |
| Bespoke Labs Data curation and RL environment recipes from ex-Google DeepMind researchers | Coding, Machine Learning | Pure-play commercial Catalog unknown | Mountain View, Menlo Park, Bangalore, San Francisco | 11-25 | ~$40M (reported) | Unknown |
| Datacurve Frontier coding data and repository RL environments via the Shipd bounty platform | Coding, RLHF | Data + environments Catalog unknown | San Francisco | 26-50 | Series A, $15M led by Chemistry (Oct 2025); $17.7M total | Unknown |