Alignment

Environments that probe and train model character: honesty under pressure, resistance to reward hacking, sycophancy, deception, and safe tool use. Rather than maximizing capability, these environments generate signal on whether models behave well when no one is watching.

5 companies · 3 cataloged artifacts · as of 2026-08-16

NameDomainsType / catalogLocationTeamFundingRaising
Andon Labs
Long-horizon autonomy benchmarks like Vending-Bench
Long Horizon, Alignment
Pure-play commercial
Catalog unknown
San Francisco1-10Seed (Y Combinator)Unknown
Gray Swan AI
Adversarial red-teaming arenas and safety evals for frontier models
Cybersecurity, Alignment
Pure-play commercial
Catalog unknown
Pittsburgh11-25~$40M (reported)Unknown
Invariant Labs
Security testing and analysis for AI agents; acquired by Snyk
Cybersecurity, Alignment
Acquired / inactive
Catalog unknown
Zurich1-10Acquired by Snyk (June 2025)Unknown
SynthLabs
Post-training research: synthetic data and scalable RL alignment
Machine Learning, Alignment, RLHF
Data + environments
Catalog unknown
San Francisco1-10Seed (M12 and First Spark Ventures, 2024)Unknown
Trajectory Labs
Alignment-focused environments for safe agent trajectories
Alignment
Pure-play commercial
Catalog unknown
Berkeley, Toronto1-10-Unknown

Available environments, datasets & benchmarks

Source-backed catalog records connected to alignment.

Full catalog →
Dataset

RLHF

Custom capability

Human preference and reward data for reinforcement learning from human feedback.

Surge · 100% completeCompany-reported
Evaluation Suite

Human Evaluation

Custom capability

Human evaluation programs for model quality, usefulness, safety, and subjective output characteristics.

Surge · 100% completeCompany-reported

Other domains