Skip to content
Boards/ReasoningStale since 2026-03-11

MMLU-Pro

12,000 harder, ten-option multiple-choice questions across 14 academic subjects.
TIGER-Lab
Gemini 3.1 Pro Preview
#1 right now
91.2%
Accuracy
100
Models ranked
Mar 11, 2026
Board published

Ranking

Best variant per model, as published. Higher is better.

  1. 1
    Gemini 3.1 Pro Preview
    as “Gemini-3.1-Pro
    91.2%
  2. 2
    Gemini 3 Pro
    as “Gemini-3-Pro(11/25)
    90.1%
  3. 3
    o1
    as “GPT-o1
    89.3%
  4. 4
    Claude Opus 4.6
    as “Claude-4.6-Opus(Thinking)
    89.1%
  5. 5
    Gemini 3 Flash
    as “Gemini-3-Flash(12/25)
    88.6%
  6. 6
    MiniMax M2.1
    as “MiniMax-M2.1
    88.0%
  7. 7
    Qwen3.5 397B A17B
    as “Qwen3.5-397B-A17B
    87.8%
  8. 8
    Seed2.0 Lite
    as “Seed2.0-Lite
    87.7%
  9. 987.5%
  10. 10
    Claude Sonnet 4.5
    as “Claude-4.5-Sonnet(Thinking)
    87.4%
  11. 1087.4%
  12. 12
    Claude Opus 4
    as “Claude-4-Opus-Thinking
    87.3%
  13. 12
    Claude Opus 4.5
    as “Claude-4.5-Opus(Thinking)
    87.3%
  14. 12
    Claude Sonnet 4.6
    as “Claude-4.6-Sonnet(Thinking)
    87.3%
  15. 15
    Hunyuan T1
    as “Hunyuan-T1
    87.2%
  16. 16
    GPT-5
    as “GPT-5(high)
    87.1%
  17. 16
    K2.5 1t A32B
    as “K2.5-1T-A32B
    87.1%
  18. 18
    Grok 4
    as “Grok-4
    87.0%
  19. 18
    Seed V1.5
    as “Seed-Thinking-v1.5
    87.0%
  20. 18
    Seed2.0 Pro
    as “Seed2.0-Pro
    87.0%
  21. 21
    Qwen3.5 122B A10B
    as “Qwen3.5-122B-A10B
    86.7%
  22. 22
    Seed1.6
    as “Seed1.6-Thinking
    86.6%
  23. 22
    Seed1.6 Base
    as “Seed1.6-Base
    86.6%
  24. 2486.4%
  25. 24
    Seed1.6 Ada
    as “Seed1.6-Ada-Thinking
    86.4%

Rank over time

History builds with every daily capture. One capture so far.

Other reasoning boards

Scores as published by TIGER-Lab on the capture date (2026-09-24). Source: TIGER-Lab results.csv on Hugging Face.

huggingface.co/spaces/TIGER-Lab/MMLU-Pro

Weekly: the AI leaderboards with a new #1, Saturday mornings.