Skip to content

SWE-Bench Pro

Long-horizon software engineering tasks from real repositories, public set, resolve rate.
Scale AI
Claude Opus 5
#1 right now
98.0%
Resolved
11
Models ranked
Sep 24, 2026
Board published

Ranking

Best variant per model, as published. Higher is better.

  1. 1
    Claude Opus 5
    as “Opus 5 (Claude Code) xhigh
    98.0%
  2. 2
    Claude Fable 5.1
    as “Fable 5.1 (Claude Code) high
    92.2%
  3. 3
    GPT-6 Astra
    as “GPT-6 Astra (Codex) high
    90.2%
  4. 4
    Claude Sonnet 5
    as “Sonnet 5 (Claude Code) xhigh
    88.2%
  5. 4
    Kimi K3
    as “Kimi-K3 (mini-swe-agent) max
    88.2%
  6. 6
    GPT-5.6 Terra
    as “GPT-5.6 Terra (Codex) xhigh
    86.3%
  7. 7
    GLM-5.3
    as “GLM-5.3 (mini-swe-agent) max
    84.3%
  8. 8
    GPT-5.6 Sol
    as “GPT-5.6 Sol (Codex) xhigh
    82.4%
  9. 9
    Gemini 3.8 Flash
    as “Gemini 3.8 Flash (mini-swe-agent) high
    58.8%
  10. 10
    Inkling
    as “Inkling (mini-swe-agent) xhigh
    56.9%
  11. 11
    Claude 4.5 Haiku
    as “Haiku 4.5 (Claude Code) xhigh
    25.5%

Rank over time

History builds with every daily capture. One capture so far.

Other coding boards

Scores as published by Scale AI on the capture date (2026-09-24). Source: scale.com leaderboard page.

scale.com/leaderboard/swe_bench_pro_public_v2

Weekly: the AI leaderboards with a new #1, Saturday mornings.