Alignment
Environments that probe and train model character: honesty under pressure, resistance to reward hacking, sycophancy, deception, and safe tool use. Rather than maximizing capability, these environments generate signal on whether models behave well when no one is watching.
6 startups
| Name | Domains | Location | Team | Funding | Raising |
|---|---|---|---|---|---|
| Andon Labs Long-horizon autonomy benchmarks like Vending-Bench | Long Horizon, Alignment | San Francisco | 1-10 | Seed (Y Combinator) | - |
| dmodel Alignment-oriented environments focused on reward quality | Machine Learning, Alignment | San Francisco | 11-25 | - | - |
| Gray Swan AI Adversarial red-teaming arenas and safety evals for frontier models | Cybersecurity, Alignment | Pittsburgh | 11-25 | ~$40M (reported) | - |
| Invariant Labs Security testing and analysis for AI agents; acquired by Snyk | Cybersecurity, Alignment | Zurich | 1-10 | Acquired by Snyk (June 2025) | No |
| SynthLabs Post-training research: synthetic data and scalable RL alignment | Machine Learning, Alignment, RLHF | San Francisco | 1-10 | Seed (M12 and First Spark Ventures, 2024) | - |
| Trajectory Labs Alignment-focused environments for safe agent trajectories | Alignment | Berkeley, Toronto | 1-10 | - | - |