SWE-Bench Pro
Long-horizon software engineering tasks from real repositories, public set, resolve rate.
Claude Opus 5
#1 right now
98.0%
Resolved
11
Models ranked
Sep 24, 2026
Board published
Ranking
Best variant per model, as published. Higher is better.
- 198.0%
Claude Opus 5
as “Opus 5 (Claude Code) xhigh” - 292.2%
Claude Fable 5.1
as “Fable 5.1 (Claude Code) high” - 390.2%
GPT-6 Astra
as “GPT-6 Astra (Codex) high” - 488.2%
Claude Sonnet 5
as “Sonnet 5 (Claude Code) xhigh” - 488.2%
Kimi K3
as “Kimi-K3 (mini-swe-agent) max” - 686.3%
GPT-5.6 Terra
as “GPT-5.6 Terra (Codex) xhigh” - 784.3%
GLM-5.3
as “GLM-5.3 (mini-swe-agent) max” - 882.4%
GPT-5.6 Sol
as “GPT-5.6 Sol (Codex) xhigh” - 958.8%
Gemini 3.8 Flash
as “Gemini 3.8 Flash (mini-swe-agent) high” - 10Inkling56.9%as “Inkling (mini-swe-agent) xhigh”
- 1125.5%
Claude 4.5 Haiku
as “Haiku 4.5 (Claude Code) xhigh”
Rank over time
History builds with every daily capture. One capture so far.
Other coding boards
Scores as published by Scale AI on the capture date (2026-09-24). Source: scale.com leaderboard page.