LiveBench
Contamination-limited questions refreshed regularly: math, reasoning, coding, data analysis, language, instruction following.
Claude Fable 5.1
#1 right now
83.4%
Global average
59
Models ranked
Jun 25, 2026
Board published
Ranking
Best variant per model, as published. Higher is better.
- 183.4%
Claude Fable 5.1
as “claude-fable-5-1-max-effort” - 283.2%
Claude Opus 5.5
as “claude-opus-5-5-max-effort” - 383.0%
Claude Fable 5
as “claude-fable-5-max-effort” - 482.2%
GPT-6 Astra
as “gpt-6-astra-max” - 581.6%
Muse Spark 1.3
as “muse-spark-1.3-xhigh” - 681.1%
DeepSeek V4.1 Flash
as “deepseek-v4.1-flash-max” - 781.0%
GPT-5.6 Sol
as “gpt-5.6-sol-max” - 880.2%
GPT-5.5
as “gpt-5.5-xhigh” - 980.1%
Claude Opus 5
as “claude-opus-5-max-effort” - 10Smaug Agentic79.5%as “smaug-agentic”
- 1179.3%
GPT-6 Sol
as “gpt-6-sol-max” - 1279.2%
Kimi K3
as “kimi-k3” - 1378.8%
Gemini 3.7 Flash
as “gemini-3.7-flash-high” - 1478.5%
Qwen 3.8 Max
as “qwen3.8-max” - 1578.0%
Grok 4.6
as “grok-4.6” - 1678.0%
GPT-5.4
as “gpt-5.4-xhigh” - 1778.0%
Muse Spark 1.2
as “muse-spark-1.2-xhigh” - 1877.9%
GPT-5.6 Terra
as “gpt-5.6-terra-max” - 1977.4%
DeepSeek V4 Pro 0813
as “deepseek-v4-pro-0813” - 19Smaug Flash77.4%as “smaug-flash”
- 2177.4%
Grok 4.7
as “grok-4.7-xhigh” - 2277.0%
Gemini 3.1 Pro Preview
as “gemini-3.1-pro-preview-high” - 23Smaug Mini76.9%as “smaug-mini”
- 2476.8%
DeepSeek V4 Flash Vision
as “deepseek-v4-flash-vision-exp” - 2576.5%
Claude Opus 4.7
as “claude-opus-4-7-xhigh-effort”
Rank over time
Rank at each month’s end, from snapshots this board published or we captured.
Other reasoning boards
AA Intelligence
#1 Claude Opus 5.5
HLE
#1 GPT-6 Astra
GPQA Diamond
#1 GPT-6 Astra
Epoch ECI
#1 GPT-6 Astra
ARC-AGI-2
#1 GPT-6 Astra
MMLU-Pro
#1 Gemini 3.1 Pro Preview
Open LLM LB
#1 none
Scores as published by LiveBench on the capture date (2026-09-24). Source: livebench.ai release tables.