Pure-play commercialCatalog available · 4

BenchFlow

Open-source benchmark hub and eval infrastructure for agents

BenchFlow builds open-source evaluation infrastructure and a hub for agent benchmarks spanning terminal, code, browser, and enterprise tasks. Its benchmarks include SkillsBench.

SkillsBench

UnknownConfidence: unknown·Last verified: Unknown·How verification works

This legacy directory profile is awaiting claim-level source migration. Existing values are retained, not upgraded to verified facts.

Websitebenchflow.ai
DomainsEnterprise, Browser, Coding
LocationSan Francisco
Team size1-10
Funding~$1M (reported)
FoundersXiangyi Li @xdotli

Capability coverage

Focus areas and technical capabilities are shown separately. Missing technical evidence remains Unknown.

Focus areas

Legacy classification; source migration pending

Catalog-evidenced technical capabilities

Unknown. No source-backed catalog capability record is available yet.

Capability catalog

Products and services supported by official company materials. Reviewed 2026-08-16.

Suggest an item
Benchmarkcompany reported

SkillsBench

Public agent-skills benchmark from BenchFlow's research catalog.

Public / Open
SkillsBench
Benchmarkcompany reported

ClawsBench

Public benchmark from BenchFlow's research catalog.

Public / Open
ClawsBench
Product capabilitycompany reported

PostTrain

Post-training product listed by BenchFlow.

Commercial
Official website
Product capabilitycompany reported

Environment Runtime

Runtime for agent environments.

Commercial
Official website

Environments & datasets

Verified public artifacts and documented private commercial inventory.

Add catalog item →
No verified catalog items recorded yet. This does not mean the company has no environments or datasets.

Sources

Structured sources have not yet been migrated for this profile.

Improve this record

Company representatives and researchers can propose sourced changes. Submissions do not directly overwrite editorial data.

Similar startups

Also listed under Enterprise, Browser, Coding.

NameDomainsType / catalogLocationTeamFundingRaising
Rise Data Labs
RL environments and tasks across multiple domains, supported by a large expert network
Computer Use, Coding, Finance, Cybersecurity, Legal, Data Labeling, Enterprise, Multi-Domain, RLHF, Custom Environments
Data + environments
3 artifacts
New York, United States11-25-Unknown
AfterQuery
Expert human data and RL environments across code, finance, and computer use
Multi-Domain, Coding, Finance
Pure-play commercial
1 artifacts
San Francisco, New York, Seattle51-100$30.5M total (reported)Unknown
Akhara
Enterprise and code RL environments
Enterprise, Coding
Pure-play commercial
Catalog unknown
San Francisco1-10-Unknown
Applied Compute
Ex-OpenAI trio applying RL to build specialist enterprise models
Enterprise, Machine Learning, Custom Environments
Pure-play commercial
Catalog unknown
San Francisco11-25$80M total at $700M valuation (Oct 2025)Yes
Bespoke Labs
Data curation and RL environment recipes from ex-Google DeepMind researchers
Coding, Machine Learning
Pure-play commercial
Catalog unknown
Mountain View, Menlo Park, Bangalore, San Francisco11-25~$40M (reported)Unknown
Collinear
Enterprise simulation, judges, and long-horizon trajectory generation
Enterprise, Long Horizon, Machine Learning, Simulation
Pure-play commercial
Catalog unknown
Mountain View, Sunnyvale11-25-Unknown