Machine Learning
Environments where agents do machine learning itself: reproducing papers, debugging training runs, tuning models, and improving benchmark scores under compute budgets. These are among the hardest long-horizon environments, and the most strategically interesting, since they measure a model's ability to improve models.
22 startups
| Name | Domains | Location | Team | Funding | Raising |
|---|---|---|---|---|---|
| Applied Compute Ex-OpenAI trio applying RL to build specialist enterprise models | Enterprise, Machine Learning, Custom Environments | San Francisco | 11-25 | $80M total at $700M valuation (Oct 2025) | Yes |
| Artificial Analysis Independent benchmarking of AI models across intelligence, speed, and price | Multi-Domain, Machine Learning | San Francisco | 11-25 | $2.6M (2024) | - |
| Bespoke Labs Data curation and RL environment recipes from ex-Google DeepMind researchers | Coding, Machine Learning | Mountain View, Menlo Park, Bangalore, San Francisco | 11-25 | ~$40M (reported) | - |
| Collinear Enterprise simulation, judges, and long-horizon trajectory generation | Enterprise, Long Horizon, Machine Learning, Simulation | Mountain View, Sunnyvale | 11-25 | - | - |
| Diffuse Labs ML and long-horizon RL environments | Machine Learning, Long Horizon | Palo Alto, San Francisco | 1-10 | - | - |
| dmodel Alignment-oriented environments focused on reward quality | Machine Learning, Alignment | San Francisco | 11-25 | - | - |
| EdotEnv Long-horizon planning environments for frontier models | Long Horizon, Machine Learning | San Francisco | 1-10 | - | - |
| Emulated Code and ML RL environments | Coding, Machine Learning | San Francisco | 1-10 | - | - |
| Epoch AI Nonprofit research institute behind FrontierMath and AI capability benchmarks | Math, Machine Learning | Remote | 11-25 | Philanthropic grants (nonprofit) | - |
| General Reasoning Open reasoning data and reward models from the ex-Meta AI reasoning lead | Finance, Long Horizon, Machine Learning | London, San Francisco | 1-10 | ~$10.9M (reported) | - |
| Genesis AI Physics simulation engine and foundation model for robotics | Robotics, Simulation, Machine Learning | San Francisco, Paris | 26-50 | Seed, $105M (Eclipse Ventures, Khosla Ventures, July 2025) | - |
| LMArena Crowdsourced model leaderboards from the Chatbot Arena team | Multi-Domain, Machine Learning | San Francisco, Berkeley | 26-50 | $100M seed (a16z, UC Investments, 2025); $150M at $1.7B valuation (Jan 2026) | - |
| Nous Research Open AI lab behind the Atropos RL environments framework | Machine Learning, Agents Infrastructure | New York | 11-25 | Series A, $50M led by Paradigm (~$1B valuation, 2025) | - |
| Osmosis Forward-deployed reinforcement learning for AI agents | Machine Learning | San Francisco | 1-10 | Seed, $7M (CRV, Audacious Ventures, YC) | - |
| Preference Model Stealth startup working on preference and reward modeling | Machine Learning, Coding | San Francisco, Toronto, Seattle | 11-25 | - | - |
| Prime Intellect Open superintelligence stack: compute, RL environments hub, and sandboxes | Machine Learning, Agents Infrastructure, Multi-Domain | San Francisco | 26-50 | Series A, $130M at $1B valuation (2026); $150M+ total | - |
| Scaled Foundations GRID: a simulation-first platform for robot learning | Robotics, Simulation, Machine Learning | Seattle | 1-10 | - | - |
| Silverstream AI Infrastructure and training data for reliable autonomous web agents | Browser, Machine Learning | San Francisco | 1-10 | Pre-seed, $1.2M (Gradient Ventures) | - |
| Snorkel Programmatic data platform expanding into expert evals and RL environments | Multi-Domain, Coding, Machine Learning, Data Labeling, RLHF | San Francisco, Redwood City, New York | 101-250 | Series D, $100M at $1.3B valuation (2025) | - |
| SynthLabs Post-training research: synthetic data and scalable RL alignment | Machine Learning, Alignment, RLHF | San Francisco | 1-10 | Seed (M12 and First Spark Ventures, 2024) | - |
| The LLM Data Company Evals and reward data tooling for LLM training | Machine Learning, Data Labeling, RLHF | - | 1-10 | - | - |
| Vmax Converts proprietary data into RL environments | Machine Learning, Custom Environments | San Francisco, New York | 1-10 | - | - |