Machine Learning

Environments where agents do machine learning itself: reproducing papers, debugging training runs, tuning models, and improving benchmark scores under compute budgets. These are among the hardest long-horizon environments, and the most strategically interesting, since they measure a model's ability to improve models.

22 startups

NameDomainsLocationTeamFundingRaising
Applied Compute
Ex-OpenAI trio applying RL to build specialist enterprise models
Enterprise, Machine Learning, Custom EnvironmentsSan Francisco11-25$80M total at $700M valuation (Oct 2025)Yes
Artificial Analysis
Independent benchmarking of AI models across intelligence, speed, and price
Multi-Domain, Machine LearningSan Francisco11-25$2.6M (2024)-
Bespoke Labs
Data curation and RL environment recipes from ex-Google DeepMind researchers
Coding, Machine LearningMountain View, Menlo Park, Bangalore, San Francisco11-25~$40M (reported)-
Collinear
Enterprise simulation, judges, and long-horizon trajectory generation
Enterprise, Long Horizon, Machine Learning, SimulationMountain View, Sunnyvale11-25--
Diffuse Labs
ML and long-horizon RL environments
Machine Learning, Long HorizonPalo Alto, San Francisco1-10--
dmodel
Alignment-oriented environments focused on reward quality
Machine Learning, AlignmentSan Francisco11-25--
EdotEnv
Long-horizon planning environments for frontier models
Long Horizon, Machine LearningSan Francisco1-10--
Emulated
Code and ML RL environments
Coding, Machine LearningSan Francisco1-10--
Epoch AI
Nonprofit research institute behind FrontierMath and AI capability benchmarks
Math, Machine LearningRemote11-25Philanthropic grants (nonprofit)-
General Reasoning
Open reasoning data and reward models from the ex-Meta AI reasoning lead
Finance, Long Horizon, Machine LearningLondon, San Francisco1-10~$10.9M (reported)-
Genesis AI
Physics simulation engine and foundation model for robotics
Robotics, Simulation, Machine LearningSan Francisco, Paris26-50Seed, $105M (Eclipse Ventures, Khosla Ventures, July 2025)-
LMArena
Crowdsourced model leaderboards from the Chatbot Arena team
Multi-Domain, Machine LearningSan Francisco, Berkeley26-50$100M seed (a16z, UC Investments, 2025); $150M at $1.7B valuation (Jan 2026)-
Nous Research
Open AI lab behind the Atropos RL environments framework
Machine Learning, Agents InfrastructureNew York11-25Series A, $50M led by Paradigm (~$1B valuation, 2025)-
Osmosis
Forward-deployed reinforcement learning for AI agents
Machine LearningSan Francisco1-10Seed, $7M (CRV, Audacious Ventures, YC)-
Preference Model
Stealth startup working on preference and reward modeling
Machine Learning, CodingSan Francisco, Toronto, Seattle11-25--
Prime Intellect
Open superintelligence stack: compute, RL environments hub, and sandboxes
Machine Learning, Agents Infrastructure, Multi-DomainSan Francisco26-50Series A, $130M at $1B valuation (2026); $150M+ total-
Scaled Foundations
GRID: a simulation-first platform for robot learning
Robotics, Simulation, Machine LearningSeattle1-10--
Silverstream AI
Infrastructure and training data for reliable autonomous web agents
Browser, Machine LearningSan Francisco1-10Pre-seed, $1.2M (Gradient Ventures)-
Snorkel
Programmatic data platform expanding into expert evals and RL environments
Multi-Domain, Coding, Machine Learning, Data Labeling, RLHFSan Francisco, Redwood City, New York101-250Series D, $100M at $1.3B valuation (2025)-
SynthLabs
Post-training research: synthetic data and scalable RL alignment
Machine Learning, Alignment, RLHFSan Francisco1-10Seed (M12 and First Spark Ventures, 2024)-
The LLM Data Company
Evals and reward data tooling for LLM training
Machine Learning, Data Labeling, RLHF-1-10--
Vmax
Converts proprietary data into RL environments
Machine Learning, Custom EnvironmentsSan Francisco, New York1-10--

Other domains