ARC-AGI-2
Abstract visual puzzles that are easy for people and hard for AI. Semi-private set, systems that are plain models only.
GPT-6 Astra
#1 right now
95.0%
Score
78
Models ranked
Sep 22, 2026
Board published
Ranking
Best variant per model, as published. Higher is better.
- 195.0%
GPT-6 Astra
as “GPT-6 Astra (Max)” - 293.3%
Claude Opus 5.5
as “Claude Opus 5.5 (High)” - 392.5%
GPT-5.6 Sol
as “GPT-5.6 Sol (Max)” - 490.4%
Claude Opus 5
as “Claude Opus 5 (Max)” - 590.0%
Claude Fable 5.1
as “Claude Fable 5.1 (Max)” - 689.2%
Claude Fable 5
as “Claude Fable 5 (Max)” - 785.0%
GPT-5.5
as “GPT-5.5 (XHigh)” - 884.6%
Gemini 3.7 Flash
as “Gemini 3.7 Flash (High)” - 984.6%
Gemini 3
as “Gemini 3 Deep Think (2/26)” - 984.6%
GPT-5.5 Pro
as “GPT-5.5 Pro (High)” - 1183.9%
GPT-5.6 Terra
as “GPT-5.6 Terra (Max)” - 1283.3%
GPT-5.4 Pro
as “GPT-5.4 Pro (XHigh)” - 1377.1%
Gemini 3.1 Pro Preview
as “Gemini 3.1 Pro (Preview)” - 14Dots3 Note76.8%as “Dots3-Note Preview (Max)”
- 1575.8%
Claude 4.7
as “Claude 4.7 (Max)” - 1674.0%
GPT-5.4
as “GPT-5.4 (XHigh)” - 1772.1%
Claude Opus 4.8
as “Claude Opus 4.8 (High)” - 1772.1%
Gemini 3.5 Flash
as “Gemini 3.5 Flash (High)” - 1969.2%
Claude Opus 4.6
as “Claude Opus 4.6 (120K, High)” - 2067.1%
Grok 4.6
as “Grok 4.6 (XHigh)” - 2165.1%
Grok 4.20
as “Grok 4.20 (Reasoning)” - 2261.4%
DeepSeek V4 Flash Vision
as “DeepSeek V4 Flash 0731 (Max)” - 2361.3%
DeepSeek V4 Pro 0813
as “DeepSeek V4 Pro 0813 (Max)” - 2460.4%
Claude Sonnet 4.6
as “Claude Sonnet 4.6 (High)” - 2560.4%
Gemini 3.6 Flash
as “Gemini 3.6 Flash (High)”
Rank over time
History builds with every daily capture. One capture so far.
Other reasoning boards
AA Intelligence
#1 Claude Opus 5.5
HLE
#1 GPT-6 Astra
GPQA Diamond
#1 GPT-6 Astra
Epoch ECI
#1 GPT-6 Astra
LiveBench
#1 Claude Fable 5.1
MMLU-Pro
#1 Gemini 3.1 Pro Preview
Open LLM LB
#1 none
Scores as published by ARC Prize on the capture date (2026-09-24). Source: arcprize.org leaderboard data files.