#29
Overall rank
4.7
Podium score
0
Board wins
Feb 17, 2026
Released
Chat
Which model people prefer in blind head-to-head conversations.
Reasoning
Hard questions with checkable answers: science, math, puzzles, expert exams.
12
MMLU-Pro24
ARC-AGI-238
Epoch ECI43
LiveBench47
GPQA Diamond
87.3% · of 100 · stale since 2026-03-11
as “Claude-4.6-Sonnet(Thinking)”
60.4% · of 78
as “Claude Sonnet 4.6 (High)”
152.3 · of 100
as “Claude Sonnet 4.6”
73.0% · of 59
as “claude-sonnet-4-6-thinking-auto-medium-effort”
87.4% · of 100
as “claude-sonnet-4-6_32K”
Coding
Fixing real repositories, building web apps, editing code.
Agents and tools
Calling functions and tools, using MCP servers, searching the web.
Vision
Understanding images, charts and documents.
More from Anthropic
Claude Opus 5.5Claude Fable 5.1Claude Opus 5Claude Sonnet 5Claude Fable 5Claude Opus 4.8Claude Opus 4.7Claude Opus 4.6
Scores as published by each board on the capture date. Model names and logos belong to their owners; logos via logo.dev.