Diana 2.0 is now available in early access.Join the waitlist

Research

Benchmarks and research that raise the bar.

Private evaluations, safety research, and frontier benchmarks that help ambitious AI teams measure what actually matters.

SEAL

Safety, Evaluation, and Alignment Lab

SEAL develops evaluation methods, red-teaming techniques, and alignment research used across frontier model programs.

Private evaluations

Custom suites that probe capability, reliability, and policy adherence without leaking to public training corpora.

Safety research

Adversarial testing and mitigation playbooks for high-stakes deployments.

Leaderboards

The standard every frontier model is measured against.

Public and private leaderboards help teams compare models on hard, realistic tasks — not synthetic demos.

SWE-Bench Pro

Raising the bar for agentic coding evaluation with production-like repositories and constraints.

Learn more

Multimodal suites

Vision, audio, and document benchmarks that stress real operational workflows.

Frequently asked questions

Yes. Research engagements can include private evaluation suites tailored to your domain, risk profile, and deployment constraints.

Partner with Scale Research

Bring evaluation rigor to the systems that matter most.