SkillsBench
Public agent-skills benchmark from BenchFlow's research catalog.
Open-source benchmark hub and eval infrastructure for agents
BenchFlow builds open-source evaluation infrastructure and a hub for agent benchmarks spanning terminal, code, browser, and enterprise tasks. Its benchmarks include SkillsBench.
SkillsBench
This legacy directory profile is awaiting claim-level source migration. Existing values are retained, not upgraded to verified facts.
| Website | benchflow.ai |
|---|---|
| Domains | Enterprise, Browser, Coding |
| Location | San Francisco |
| Team size | 1-10 |
| Funding | ~$1M (reported) |
| Founders | Xiangyi Li @xdotli |
Focus areas and technical capabilities are shown separately. Missing technical evidence remains Unknown.
Unknown. No source-backed catalog capability record is available yet.
Products and services supported by official company materials. Reviewed 2026-08-16.
Public agent-skills benchmark from BenchFlow's research catalog.
Public benchmark from BenchFlow's research catalog.
Post-training product listed by BenchFlow.
Runtime for agent environments.
Verified public artifacts and documented private commercial inventory.
Structured sources have not yet been migrated for this profile.
Company representatives and researchers can propose sourced changes. Submissions do not directly overwrite editorial data.
Also listed under Enterprise, Browser, Coding.
| Name | Domains | Type / catalog | Location | Team | Funding | Raising |
|---|---|---|---|---|---|---|
| Rise Data Labs RL environments and tasks across multiple domains, supported by a large expert network | Computer Use, Coding, Finance, Cybersecurity, Legal, Data Labeling, Enterprise, Multi-Domain, RLHF, Custom Environments | Data + environments 3 artifacts | New York, United States | 11-25 | - | Unknown |
| AfterQuery Expert human data and RL environments across code, finance, and computer use | Multi-Domain, Coding, Finance | Pure-play commercial 1 artifacts | San Francisco, New York, Seattle | 51-100 | $30.5M total (reported) | Unknown |
| Akhara Enterprise and code RL environments | Enterprise, Coding | Pure-play commercial Catalog unknown | San Francisco | 1-10 | - | Unknown |
| Applied Compute Ex-OpenAI trio applying RL to build specialist enterprise models | Enterprise, Machine Learning, Custom Environments | Pure-play commercial Catalog unknown | San Francisco | 11-25 | $80M total at $700M valuation (Oct 2025) | Yes |
| Bespoke Labs Data curation and RL environment recipes from ex-Google DeepMind researchers | Coding, Machine Learning | Pure-play commercial Catalog unknown | Mountain View, Menlo Park, Bangalore, San Francisco | 11-25 | ~$40M (reported) | Unknown |
| Collinear Enterprise simulation, judges, and long-horizon trajectory generation | Enterprise, Long Horizon, Machine Learning, Simulation | Pure-play commercial Catalog unknown | Mountain View, Sunnyvale | 11-25 | - | Unknown |