How the leader of leaders is scored
- 1. Capture
Every board is fetched daily at 04:30 UTC from its own page, data file or API and parsed without a model. If a parser breaks, the model reads that board’s page instead and those rows carry an Unverified badge.
- 2. One model, one entry
Variants collapse to the model: effort levels (high, xhigh, max), thinking modes, dated snapshots and agent harnesses. The best variant on a board stands for the model there. The exact name each board published is kept next to every entry.
- 3. Which boards count
Language-model quality boards (chat, reasoning, coding, agents and tools, vision) that published in the last 180 days. Image, video, speech, embeddings, speed, price and usage get their own category leaders.
- 4. Points
On each counted board: #1 earns 10 points, #2 earns 9, down to 1 point for #10. Podium score = points earned / (10 × boards counted) × 100. Ties break on #1 finishes, then the median rank across boards that list the model.
Every board and whether it counts
15 boards count toward the leader of leaders today.
| Board | Category | Metric | Published | Counts |
|---|---|---|---|---|
lmarena.ai leaderboard page data | Chat | Arena rating | Sep 13, 2026 | Yes |
scale.com leaderboard page | Chat | Score | Sep 24, 2026 | Yes |
artificialanalysis.ai models page data | Reasoning | Index | Sep 24, 2026 | Yes |
scale.com leaderboard page | Reasoning | Accuracy | Sep 24, 2026 | Yes |
epoch.ai benchmark data (CC BY) | Reasoning | Accuracy | Sep 2, 2026 | Yes |
epoch.ai eci_scores.csv (CC BY) | Reasoning | ECI | Sep 24, 2026 | Yes |
arcprize.org leaderboard data files | Reasoning | Score | Sep 22, 2026 | Yes |
livebench.ai release tables | Reasoning | Global average | Jun 25, 2026 | Yes |
TIGER-Lab results.csv on Hugging Face | Reasoning | Accuracy | Stale since 2026-03-11 | Not fresh |
lmarena.ai leaderboard page data | Coding | Arena rating | Sep 23, 2026 | Yes |
scale.com leaderboard page | Coding | Resolved | Sep 24, 2026 | Yes |
swe-bench.github.io leaderboards.json | Coding | Resolved | Stale since 2026-02-26 | Not fresh |
Aider polyglot_leaderboard.yml on GitHub | Coding | Pass rate | Stale since 2025-10-03 | Not fresh |
gorilla.cs.berkeley.edu data_overall.csv | Agents and tools | Overall accuracy | Apr 12, 2026 | Yes |
scale.com leaderboard page | Agents and tools | Pass rate | Sep 23, 2026 | Yes |
lmarena.ai leaderboard page data | Agents and tools | Arena rating | Aug 24, 2026 | Yes |
lmarena.ai leaderboard page data | Vision | Arena rating | Sep 13, 2026 | Yes |
lmarena.ai leaderboard page data | Vision | Arena rating | Sep 13, 2026 | Yes |
lmarena.ai leaderboard page data | Image | Arena rating | Sep 21, 2026 | Category only |
lmarena.ai leaderboard page data | Image | Arena rating | Sep 21, 2026 | Category only |
lmarena.ai leaderboard page data | Video | Arena rating | Sep 21, 2026 | Category only |
lmarena.ai leaderboard page data | Video | Arena rating | Sep 21, 2026 | Category only |
artificialanalysis.ai speech arena page data | Speech | Arena rating | Sep 24, 2026 | Category only |
hf-audio results CSV on Hugging Face | Speech | Average WER (lower is better) | Sep 22, 2026 | Category only |
MTEB leaderboard backend API | Embeddings | Mean task score | Sep 22, 2026 | Category only |
artificialanalysis.ai models page data | Speed | Output speed | Sep 24, 2026 | Category only |
artificialanalysis.ai models page data | Price | Blended price (lower is better) | Sep 24, 2026 | Category only |
openrouter.ai rankings page data | Usage | Weekly tokens | Sep 23, 2026 | Category only |
retired, not captured | Reasoning | Average | Retired | Category only |
What this is not
A board is one team’s method on one day. Leaders Board does not run evaluations and does not adjust anyone’s numbers; it shows where the boards agree. Vendors: OpenAI, Anthropic, Google and others appear by the name each board uses, linked to their own pages.
Scores as published by each board on the capture date. Model names and logos belong to their owners; logos via logo.dev.