#7
Overall rank
22.7
Podium score
0
Board wins
Feb 5, 2026
Released
Chat
Which model people prefer in blind head-to-head conversations.
Reasoning
Hard questions with checkable answers: science, math, puzzles, expert exams.
4
MMLU-Pro10
HLE19
ARC-AGI-225
Epoch ECI30
GPQA Diamond37
LiveBench
89.1% · of 100 · stale since 2026-03-11
as “Claude-4.6-Opus(Thinking)”
34.4% · of 42
as “claude-opus-4-6-thinking-max”
69.2% · of 78
as “Claude Opus 4.6 (120K, High)”
155.4 · of 100
as “Claude Opus 4.6”
90.5% · of 100
as “claude-opus-4-6_32K”
74.5% · of 59
as “claude-opus-4-6-thinking-auto-high-effort”
Coding
Fixing real repositories, building web apps, editing code.
Agents and tools
Calling functions and tools, using MCP servers, searching the web.
Vision
Understanding images, charts and documents.
More from Anthropic
Claude Opus 5.5Claude Fable 5.1Claude Opus 5Claude Sonnet 5Claude Fable 5Claude Opus 4.8Claude Opus 4.7Claude Sonnet 4.6
Scores as published by each board on the capture date. Model names and logos belong to their owners; logos via logo.dev.