Scale SEAL MultiChallenge
Multi-turn conversations that test instruction retention, inference memory and self-coherence.
Muse Spark
#1 right now
75.5%
Score
29
Models ranked
Sep 24, 2026
Board published
Ranking
Best variant per model, as published. Higher is better.
- 175.5%
- 275.3%
- 371.4%
Gemini 3.1 Pro Preview
as “gemini-3.1-pro-preview” - 469.2%
GPT-5.4 Pro
as “gpt-5.4-pro-2026-03-05” - 565.7%
Gemini 3 Pro
as “gemini-3-pro-preview” - 663.4%
GPT-5.1
as “gpt-5.1-2025-11-13-thinking” - 763.2%
GPT-5
as “gpt-5-thinking” - 862.4%
o3-pro
as “o3-pro-2025-06-10-reasoning-high” - 961.4%
Kimi K2.5
as “kimi-k2.5” - 1060.6%
Gemini 3.1 Flash-Lite
as “gemini-3.1-flash-lite-preview” - 1159.0%
GPT-5 mini
as “gpt-5-mini-thinking” - 1259.0%
Claude Opus 4.5
as “claude-opus-4-5-20251101-thinking” - 1358.6%
Claude Opus 4
as “claude-4-opus-thinking” - 1457.2%
Claude Opus 4.1
as “claude-opus-4-1-20250805-thinking” - 1557.1%
Claude Sonnet 4
as “claude-4-sonnet-thinking” - 1656.6%
o3
as “o3-2025-04-16-reasoning-high” - 1756.0%
Claude Opus 4.6
as “claude-opus-4-6 (Non-Thinking)” - 1855.4%
Kimi K2 Thinking
as “kimi-k2-thinking” - 1955.3%
Claude Sonnet 4.5
as “claude-sonnet-4-5-20250929-thinking” - 2053.6%
Gemini 2.5 Pro (Jun 2025)
as “gemini-2-5-pro” - 2151.6%
Claude 3 7 Sonnet
as “claude-3-7-sonnet-thinking” - 2251.2%
GPT-5.1 Instant
as “gpt-5.1-2025-11-13-instant” - 2350.5%
Claude 4.5 Haiku
as “claude-haiku-4-5-20251001-thinking” - 2446.1%
DeepSeek V3P1
as “deepseek-v3p1” - 2545.3%
Rank over time
History builds with every daily capture. One capture so far.
Other chat boards
Scores as published by Scale AI on the capture date (2026-09-24). Source: scale.com leaderboard page.