#25
Overall rank
6.0
Podium score
0
Board wins
Sep 29, 2025
Released
Chat
Which model people prefer in blind head-to-head conversations.
Reasoning
Hard questions with checkable answers: science, math, puzzles, expert exams.
10
MMLU-Pro21
HLE42
ARC-AGI-263
Epoch ECI67
GPQA Diamond
87.4% · of 100 · stale since 2026-03-11
as “Claude-4.5-Sonnet(Thinking)”
13.7% · of 42
as “claude-sonnet-4-5-20250929-thinking”
13.6% · of 78
as “Claude Sonnet 4.5 (Thinking 32K)”
146.8 · of 100
as “Claude Sonnet 4.5”
82.3% · of 100
as “claude-sonnet-4-5-20250929_59K”
Coding
Fixing real repositories, building web apps, editing code.
Agents and tools
Calling functions and tools, using MCP servers, searching the web.
Vision
Understanding images, charts and documents.
More from Anthropic
Claude Opus 5.5Claude Fable 5.1Claude Opus 5Claude Sonnet 5Claude Fable 5Claude Opus 4.8Claude Opus 4.7Claude Sonnet 4.6
Scores as published by each board on the capture date. Model names and logos belong to their owners; logos via logo.dev.