Benchmarks

The tests we keep score with

Each was “impossible” until it wasn’t.

ARC-AGI-3

Interactive mini-games with no instructions — the agent must explore, infer the rules, and learn across levels.

~30%humans 100%

Claude Opus 5, Jul 2026

Moves past static puzzles to agentic, interactive learning — the closest test yet to "figure out a brand-new world from scratch," like a human would.

Opus 5 nearly 4×’d the previous record (7.8%) in one release — but 19 of 25 games remain unbeaten.

ARC Prize leaderboard

Humanity's Last Exam

2,500 expert-level questions across 100+ subjects, written to be un-Googleable.

~65%no single-human baseline

Claude Opus 5, with tools · Jul 2026

A fresh frontier-of-human-knowledge yardstick, built once MMLU and GPQA started saturating.

Single digits at launch (Jan 2025); best no-tools text-only runs sit in the low 50s.

Scale SEAL leaderboard

FrontierMath

Hundreds of original, unpublished research-grade math problems, auto-graded.

~88%expert teams ~19%

GPT-5.5 Pro / Claude Fable 5, v2 · 2026

Probes genuine mathematical reasoning, not recall — problems take experts hours to weeks.

FrontierMath v2 (Jun 2026) allows tools; the older no-tools set topped ~52%.

Epoch AI

METR Time Horizon

The length of task (in human time) a model can finish autonomously with 50% reliability.

~12 hrsdoubling ~every 4 months

Claude Opus 4.6 · 2026

Measures long-horizon autonomy — arguably the capability most missing from today’s AI.

Up from ~1 hr in early 2025. Mythos Preview leads at ~17 hrs — past METR’s ~16 hr reliability limit.

METR

GPQA Diamond

198 'Google-proof' graduate-level science questions (bio, physics, chemistry).

~95%PhD experts ~70%

GPT-5.6 Ultra, Jul 2026

Tests real expert reasoning vs. web lookup — a standard frontier-capability headline.

Saturated — frontier models cluster in the mid-90s.

Epoch AI

SWE-bench Pro

Resolve real, contamination-resistant GitHub issues so the hidden tests pass.

~80%no formal human baseline

Claude Mythos 5 / Fable 5, Jul 2026 (lab-reported)

The current bar for autonomous, agentic software engineering in an unfamiliar codebase.

Lab-reported on own scaffolds; Scale's standardized public leaderboard still tops out near 59%.

Scale SEAL leaderboard

ARC-AGI-1

The original 2019 abstract-reasoning test: solve novel grid puzzles from a few examples.

98%human panel 98%

Gemini 3.1 Pro, Feb 2026

For years the iconic 'easy for humans, hard for AI' benchmark.

Essentially solved (o3 cracked it in Dec 2024).

ARC Prize leaderboard

MMLU

57-subject multiple-choice exam — the old standard for broad knowledge.

~92%domain experts ~89.8%

frontier labs (self-reported), low-90s

Was the canonical proxy for general, cross-domain competence.

Saturated and retired; standardized boards top ~88%.

Epoch AI