Skip to content

How the leader of leaders is scored

Every step, so you can check it or disagree with it.
  1. 1. Capture

    Every board is fetched daily at 04:30 UTC from its own page, data file or API and parsed without a model. If a parser breaks, the model reads that board’s page instead and those rows carry an Unverified badge.

  2. 2. One model, one entry

    Variants collapse to the model: effort levels (high, xhigh, max), thinking modes, dated snapshots and agent harnesses. The best variant on a board stands for the model there. The exact name each board published is kept next to every entry.

  3. 3. Which boards count

    Language-model quality boards (chat, reasoning, coding, agents and tools, vision) that published in the last 180 days. Image, video, speech, embeddings, speed, price and usage get their own category leaders.

  4. 4. Points

    On each counted board: #1 earns 10 points, #2 earns 9, down to 1 point for #10. Podium score = points earned / (10 × boards counted) × 100. Ties break on #1 finishes, then the median rank across boards that list the model.

101
92
83
74
65
56
47
38
29
110

Every board and whether it counts

15 boards count toward the leader of leaders today.

BoardCategoryMetricPublishedCounts
LMArena Text Arena
lmarena.ai leaderboard page data
ChatArena ratingSep 13, 2026 Yes
Scale SEAL MultiChallenge
scale.com leaderboard page
ChatScoreSep 24, 2026 Yes
Artificial Analysis Intelligence Index
artificialanalysis.ai models page data
ReasoningIndexSep 24, 2026 Yes
Humanity's Last Exam
scale.com leaderboard page
ReasoningAccuracySep 24, 2026 Yes
GPQA Diamond (Epoch AI runs)
epoch.ai benchmark data (CC BY)
ReasoningAccuracySep 2, 2026 Yes
Epoch Capabilities Index
epoch.ai eci_scores.csv (CC BY)
ReasoningECISep 24, 2026 Yes
ARC-AGI-2
arcprize.org leaderboard data files
ReasoningScoreSep 22, 2026 Yes
LiveBench
livebench.ai release tables
ReasoningGlobal averageJun 25, 2026 Yes
MMLU-Pro
TIGER-Lab results.csv on Hugging Face
ReasoningAccuracyStale since 2026-03-11Not fresh
LMArena WebDev Arena
lmarena.ai leaderboard page data
CodingArena ratingSep 23, 2026 Yes
SWE-Bench Pro
scale.com leaderboard page
CodingResolvedSep 24, 2026 Yes
SWE-bench Verified
swe-bench.github.io leaderboards.json
CodingResolvedStale since 2026-02-26Not fresh
Aider Polyglot
Aider polyglot_leaderboard.yml on GitHub
CodingPass rateStale since 2025-10-03Not fresh
Berkeley Function Calling Leaderboard
gorilla.cs.berkeley.edu data_overall.csv
Agents and toolsOverall accuracyApr 12, 2026 Yes
MCP Atlas
scale.com leaderboard page
Agents and toolsPass rateSep 23, 2026 Yes
LMArena Search Arena
lmarena.ai leaderboard page data
Agents and toolsArena ratingAug 24, 2026 Yes
LMArena Vision Arena
lmarena.ai leaderboard page data
VisionArena ratingSep 13, 2026 Yes
LMArena Document Arena
lmarena.ai leaderboard page data
VisionArena ratingSep 13, 2026 Yes
LMArena Text-to-Image
lmarena.ai leaderboard page data
ImageArena ratingSep 21, 2026Category only
LMArena Image Edit
lmarena.ai leaderboard page data
ImageArena ratingSep 21, 2026Category only
LMArena Text-to-Video
lmarena.ai leaderboard page data
VideoArena ratingSep 21, 2026Category only
LMArena Image-to-Video
lmarena.ai leaderboard page data
VideoArena ratingSep 21, 2026Category only
Artificial Analysis Speech Arena
artificialanalysis.ai speech arena page data
SpeechArena ratingSep 24, 2026Category only
Open ASR Leaderboard
hf-audio results CSV on Hugging Face
SpeechAverage WER (lower is better)Sep 22, 2026Category only
MTEB Multilingual v2
MTEB leaderboard backend API
EmbeddingsMean task scoreSep 22, 2026Category only
Artificial Analysis Output Speed
artificialanalysis.ai models page data
SpeedOutput speedSep 24, 2026Category only
Artificial Analysis Price
artificialanalysis.ai models page data
PriceBlended price (lower is better)Sep 24, 2026Category only
OpenRouter weekly usage
openrouter.ai rankings page data
UsageWeekly tokensSep 23, 2026Category only
Hugging Face Open LLM Leaderboard
retired, not captured
ReasoningAverageRetiredCategory only

What this is not

A board is one team’s method on one day. Leaders Board does not run evaluations and does not adjust anyone’s numbers; it shows where the boards agree. Vendors: OpenAI, Anthropic, Google and others appear by the name each board uses, linked to their own pages.

Scores as published by each board on the capture date. Model names and logos belong to their owners; logos via logo.dev.

All boards

Weekly: the AI leaderboards with a new #1, Saturday mornings.