BenchFlow
Open-source benchmark hub and eval infrastructure for agents
BenchFlow builds open-source evaluation infrastructure and a hub for agent benchmarks spanning terminal, code, browser, and enterprise tasks. Its benchmarks include SkillsBench.
SkillsBench
| Website | benchflow.ai |
|---|---|
| Domains | Enterprise, Browser, Coding |
| Location | San Francisco |
| Team size | 1-10 |
| Funding | ~$1M (reported) |
| Founders | Xiangyi Li @xdotli |
Similar startups
Also listed under Enterprise, Browser, Coding.
| Name | Domains | Location | Team | Funding | Raising |
|---|---|---|---|---|---|
| Rise Data Labs US-based expert human data and custom RL environments for enterprise AI | Computer Use, Coding, Finance, Cybersecurity, Legal, Data Labeling, Enterprise, Multi-Domain, RLHF, Custom Environments | United States | - | - | Yes |
| AfterQuery Expert human data and RL environments across code, finance, and computer use | Multi-Domain, Coding, Finance | San Francisco, New York, Seattle | 51-100 | $30.5M total (reported) | - |
| AIChamp Custom enterprise-workflow RL environments with expert grading | Enterprise, Long Horizon, Custom Environments | San Francisco | 1-10 | - | - |
| Akhara Enterprise and code RL environments | Enterprise, Coding | San Francisco | 1-10 | - | - |
| Anchor Browser Reliable browser automation platform for agentic AI | Browser, Agents Infrastructure | Tel Aviv | 1-10 | Seed, $6M (Blumberg Capital, Gradient Ventures, Nov 2025) | - |
| Applied Compute Ex-OpenAI trio applying RL to build specialist enterprise models | Enterprise, Machine Learning, Custom Environments | San Francisco | 11-25 | $80M total at $700M valuation (Oct 2025) | Yes |