Humanity's Last Exam
2,500 expert-written questions across dozens of subjects, built to be the last closed-ended academic exam.
GPT-6 Astra
#1 right now
54.8%
Accuracy
42
Models ranked
Sep 24, 2026
Board published
Ranking
Best variant per model, as published. Higher is better.
- 154.8%
GPT-6 Astra
as “GPT 6 Astra” - 246.5%
Claude Fable 5.1
as “Fable 5.1 (xhigh)” - 346.4%
Gemini 3.1 Pro Preview
as “gemini-3.1-pro-preview (thinking high)” - 444.5%
- 544.3%
GPT-5.4 Pro
as “gpt-5.4-pro-2026-03-05” - 640.6%
- 737.5%
Gemini 3 Pro
as “gemini-3-pro-preview” - 836.2%
GPT-5.4
as “gpt-5.4-2026-03-05 (xhigh thinking)” - 936.2%
Claude Opus 4.7
as “claude-opus-4-7” - 1034.4%
Claude Opus 4.6
as “claude-opus-4-6-thinking-max” - 1131.6%
GPT-5 Pro
as “gpt-5-pro-2025-10-06” - 1227.8%
GPT-5.2
as “gpt-5.2-2025-12-11” - 1325.3%
GPT-5
as “gpt-5-2025-08-07” - 1425.2%
Claude Opus 4.5
as “claude-opus-4-5-20251101-thinking” - 1524.4%
Kimi K2.5
as “kimi-k2.5” - 1623.7%
GPT-5.1
as “gpt-5.1-thinking” - 1721.6%
Gemini 2.5 Pro (Jun 2025)
as “gemini-2.5-pro-preview-06-05” - 1820.3%
o3
as “o3 (high) (April 2025)” - 1919.4%
GPT-5 mini
as “gpt-5-mini-2025-08-07” - 2018.1%
o4-mini
as “o4-mini (high) (April 2025)” - 2113.7%
Claude Sonnet 4.5
as “claude-sonnet-4-5-20250929-thinking” - 2212.1%
Gemini 2.5 Flash (Sep 2025)
as “Gemini 2.5 Flash (April 2025)” - 2311.5%
Claude Opus 4.1
as “claude-opus-4-1-20250805-thinking” - 2410.7%
Claude Opus 4
as “Claude Opus 4 (Thinking)” - 258.6%
Gemini 3.1 Flash-Lite
as “gemini-3.1-flash-lite-preview”
Rank over time
History builds with every daily capture. One capture so far.
Other reasoning boards
AA Intelligence
#1 Claude Opus 5.5
GPQA Diamond
#1 GPT-6 Astra
Epoch ECI
#1 GPT-6 Astra
ARC-AGI-2
#1 GPT-6 Astra
LiveBench
#1 Claude Fable 5.1
MMLU-Pro
#1 Gemini 3.1 Pro Preview
Open LLM LB
#1 none
Scores as published by Scale AI and CAIS on the capture date (2026-09-24). Source: scale.com leaderboard page.