Benchmarks
The tests we keep score with
Each was “impossible” until it wasn’t.
ARC-AGI-3
Interactive mini-games with no instructions — the agent must explore, infer the rules, and learn across levels.
~30%humans 100%
Claude Opus 5, Jul 2026
Moves past static puzzles to agentic, interactive learning — the closest test yet to "figure out a brand-new world from scratch," like a human would.
Opus 5 nearly 4×’d the previous record (7.8%) in one release — but 19 of 25 games remain unbeaten.
ARC Prize leaderboardHumanity's Last Exam
2,500 expert-level questions across 100+ subjects, written to be un-Googleable.
~65%no single-human baseline
Claude Opus 5, with tools · Jul 2026
A fresh frontier-of-human-knowledge yardstick, built once MMLU and GPQA started saturating.
Single digits at launch (Jan 2025); best no-tools text-only runs sit in the low 50s.
Scale SEAL leaderboardFrontierMath
Hundreds of original, unpublished research-grade math problems, auto-graded.
~88%expert teams ~19%
GPT-5.5 Pro / Claude Fable 5, v2 · 2026
Probes genuine mathematical reasoning, not recall — problems take experts hours to weeks.
FrontierMath v2 (Jun 2026) allows tools; the older no-tools set topped ~52%.
Epoch AIMETR Time Horizon
The length of task (in human time) a model can finish autonomously with 50% reliability.
~12 hrsdoubling ~every 4 months
Claude Opus 4.6 · 2026
Measures long-horizon autonomy — arguably the capability most missing from today’s AI.
Up from ~1 hr in early 2025. Mythos Preview leads at ~17 hrs — past METR’s ~16 hr reliability limit.
METRGPQA Diamond
198 'Google-proof' graduate-level science questions (bio, physics, chemistry).
~95%PhD experts ~70%
GPT-5.6 Ultra, Jul 2026
Tests real expert reasoning vs. web lookup — a standard frontier-capability headline.
Saturated — frontier models cluster in the mid-90s.
Epoch AISWE-bench Pro
Resolve real, contamination-resistant GitHub issues so the hidden tests pass.
~80%no formal human baseline
Claude Mythos 5 / Fable 5, Jul 2026 (lab-reported)
The current bar for autonomous, agentic software engineering in an unfamiliar codebase.
Lab-reported on own scaffolds; Scale's standardized public leaderboard still tops out near 59%.
Scale SEAL leaderboardARC-AGI-1
The original 2019 abstract-reasoning test: solve novel grid puzzles from a few examples.
98%human panel 98%
Gemini 3.1 Pro, Feb 2026
For years the iconic 'easy for humans, hard for AI' benchmark.
Essentially solved (o3 cracked it in Dec 2024).
ARC Prize leaderboardMMLU
57-subject multiple-choice exam — the old standard for broad knowledge.
~92%domain experts ~89.8%
frontier labs (self-reported), low-90s
Was the canonical proxy for general, cross-domain competence.
Saturated and retired; standardized boards top ~88%.
Epoch AI