Aider Polyglot
225 hard Exercism exercises in C++, Go, Java, JavaScript, Python and Rust, edited through Aider.
GPT-5
#1 right now
88.0%
Pass rate
50
Models ranked
Oct 3, 2025
Board published
Ranking
Best variant per model, as published. Higher is better.
- 188.0%
GPT-5
as “gpt-5 (high)” - 284.9%
o3-pro
as “o3-pro (high)” - 383.1%
Gemini 2.5 Pro (Jun 2025)
as “gemini-2.5-pro-preview-06-05 (32k think)” - 481.3%
o3
as “o3 (high)” - 579.6%
Grok 4
as “grok-4 (high)” - 678.2%
o3 GPT 4.1
as “o3 (high) + gpt-4.1” - 776.9%
Gemini 2.5 Pro 05 06
as “Gemini 2.5 Pro Preview 05-06” - 874.2%
DeepSeek-V3.2
as “DeepSeek-V3.2-Exp (Reasoner)” - 972.9%
Gemini 2.5 Pro 03 25
as “Gemini 2.5 Pro Preview 03-25” - 1072.0%
Claude Opus 4
as “claude-opus-4-20250514 (32k thinking)” - 1072.0%
o4-mini
as “o4-mini (high)” - 1271.4%
DeepSeek-R1 (May 2025)
as “DeepSeek R1 (0528)” - 1364.9%
Claude 3 7 Sonnet
as “claude-3-7-sonnet-20250219 (32k thinking tokens)” - 1464.0%
DeepSeek R1 Claude 3 5 Sonnet
as “DeepSeek R1 + claude-3-5-sonnet-20241022” - 1561.7%
o1
as “o1-2024-12-17 (high)” - 1661.3%
Claude Sonnet 4
as “claude-sonnet-4-20250514 (32k thinking)” - 1760.4%
o3 Mini
as “o3-mini (high)” - 1859.6%
Qwen3 235b A22B Diff No Alibaba
as “Qwen3 235B A22B diff, no think, Alibaba API” - 1959.1%
Kimi K2 Thinking
as “Kimi K2” - 2055.1%
DeepSeek V3
as “DeepSeek V3 (0324)” - 2055.1%
Gemini 2.5 Flash (Sep 2025)
as “gemini-2.5-flash-preview-05-20 (24k think)” - 2254.7%
- 2353.3%
Grok 3
as “Grok 3 Beta” - 2452.9%
- 2552.4%
GPT 4.1
as “gpt-4.1”
Rank over time
History builds with every daily capture. One capture so far.
Other coding boards
Scores as published by Aider on the capture date (2026-09-24). Source: Aider polyglot_leaderboard.yml on GitHub.