EnvIndex catalog

Environments, datasets & benchmarks

Canonical market-intelligence pages for the artifacts vendors and research teams actually offer. Public availability, commercial access, and verification are tracked separately.

21
cataloged artifacts
8
public samples
4
private commercial
16
with verifier details

21 of 21 artifacts

Dataset

RLHF

Custom capability

Human preference and reward data for reinforcement learning from human feedback.

Surge · 100% completeCompany-reported
Evaluation Suite

Human Evaluation

Custom capability

Human evaluation programs for model quality, usefulness, safety, and subjective output characteristics.

Surge · 100% completeCompany-reported
Custom capability

Language and culturally grounded training data across more than 70 reported languages.

Surge · 100% completeCompany-reported
Custom capability

Training and evaluation data spanning text, images, audio, and video.

Surge · 100% completeCompany-reported
Benchmark

CoreCraft

Public sample

A simulated enterprise environment where agents complete customer-support and operational tasks inside a fictional PC retailer.

Surge · 100% completeCompany-reported
Environment Bundle

Finance Tasks

Private commercial

A private commercial catalog of finance tasks for RL training and evaluation.

Rise Data Labs · 100% completeVendor submitted
Benchmark

GBA Eval

Public sample

A long-horizon coding eval in which agents build a Game Boy Advance emulator in Rust within 24 hours.

Mechanize · 100% completeCompany-reported
Benchmark

SWE-Bench Pro

Public sample

A long-horizon software-engineering benchmark built from public and proprietary repositories.

Scale · 100% completeConfirmed
Benchmark

App-Bench

Public sample

A benchmark for evaluating AI systems that generate modern web applications.

AfterQuery · 100% completeCompany-reported
Benchmark

APEX

Public sample

An expert-authored benchmark of economically valuable tasks in finance, consulting, law, and medicine.

Mercor · 100% completeConfirmed