RL Environment
RL Environments and Agents
Custom RL environments and verifier design for training and evaluating agentic models.
Surge · 100% completeCompany-reported
Human preference and reward data for reinforcement learning from human feedback.
Surge reports generating preference and reward data intended to capture nuanced judgments for model alignment and post-training.
Only evidenced capabilities are shown as Yes. Missing evidence remains Unknown.
Custom RL environments and verifier design for training and evaluating agentic models.
Custom rubric and verifier design for scoring complex model and agent behavior.
Expert demonstrations for bootstrapping model capabilities, including computer and browser use.