Vending-Bench 2
Long-horizon benchmark built around running a simulated vending-machine business.
Long-horizon autonomy benchmarks like Vending-Bench
Andon Labs (YC, formerly Vectorview) builds benchmarks and evaluations for AI agents' long-horizon coherence and safety, including Vending-Bench, Butter-Bench, and Blueprint-Bench, and deploys agents into real-world businesses. It partnered with Anthropic on Project Vend, letting Claude run a real office vending machine.
Anthropic's Project Vend collaboration; Vending-Bench
This legacy directory profile is awaiting claim-level source migration. Existing values are retained, not upgraded to verified facts.
| Website | andonlabs.com |
|---|---|
| Domains | Long Horizon, Alignment |
| Location | San Francisco |
| Team size | 1-10 |
| Funding | Seed (Y Combinator) |
| Founders | Lukas Petersson, Axel Backlund |
Focus areas and technical capabilities are shown separately. Missing technical evidence remains Unknown.
Unknown. No source-backed catalog capability record is available yet.
Products and services supported by official company materials. Reviewed 2026-08-16.
Long-horizon benchmark built around running a simulated vending-machine business.
Benchmark listed in Andon Labs' public evaluation catalog.
Benchmark listed in Andon Labs' public evaluation catalog.
Benchmark listed in Andon Labs' public evaluation catalog.
Verified public artifacts and documented private commercial inventory.
Structured sources have not yet been migrated for this profile.
Company representatives and researchers can propose sourced changes. Submissions do not directly overwrite editorial data.
Also listed under Long Horizon, Alignment.
| Name | Domains | Type / catalog | Location | Team | Funding | Raising |
|---|---|---|---|---|---|---|
| Rise Data Labs RL environments and tasks across multiple domains, supported by a large expert network | Computer Use, Coding, Finance, Cybersecurity, Legal, Data Labeling, Enterprise, Multi-Domain, RLHF, Custom Environments | Data + environments 3 artifacts | New York, United States | 11-25 | - | Unknown |
| Andromede Programmatic generation of long-horizon RL environments | Long Horizon | Pure-play commercial Catalog unknown | Lausanne | 1-10 | - | Unknown |
| Anthromind Medical and long-horizon environments and expert data | Medical, Long Horizon, Data Labeling | Data + environments Catalog unknown | San Francisco | 1-10 | - | Unknown |
| ARIMLABS Security and long-horizon environments for agentic AI | Cybersecurity, Long Horizon | Pure-play commercial Catalog unknown | Warsaw | 11-25 | - | Unknown |
| Collinear Enterprise simulation, judges, and long-horizon trajectory generation | Enterprise, Long Horizon, Machine Learning, Simulation | Pure-play commercial Catalog unknown | Mountain View, Sunnyvale | 11-25 | - | Unknown |
| Diffuse Labs ML and long-horizon RL environments | Machine Learning, Long Horizon | Pure-play commercial Catalog unknown | Palo Alto, San Francisco | 1-10 | - | Unknown |