AI judge rankings

Provisional automated evaluations, kept separate from human votes.

reasoning enabled · non-reasoning = explicitly disabled · reasoning ? = reasoning mode unknown

Claude Fable 5

Partial — not a final ranking

98 validated duel verdicts

Results from Claude Fable 5
Rank Model Wins Losses Ties Comparisons
1 Qwen 3.7 Plus 💡 15 6 4 25
2 Qwen 3.7 Max 💡 13 7 5 25
3 MiMo V2.5 Pro 💡 12 8 4 24
4 MiniMax M3 💡 10 2 2 14
5 DeepSeek V4 Pro 💡 9 7 6 22
6 DeepSeek V4 Flash 💡 8 14 3 25
7 MiMo V2.5 💡 7 15 2 24
8 GLM 5.2 💡 6 14 4 24
9 Qwen 3.6 27B Q4_K_M 💡 2 1 1 4
10 Gemma 4 26B-A4B Instruct Q8_0 💡 0 1 0 1
11 Gemma 4 31B Instruct Q8_0 💡 0 1 0 1
12 Qwen 3.6 35B-A3B Q4_K_M 💡 0 1 0 1
13 Gemma 4 26B-A4B Instruct QAT 💡 0 2 0 2
14 Gemma 4 31B Instruct QAT 💡 0 3 1 4

Claude Opus 4 8

Partial — not a final ranking

435 validated duel verdicts

Results from Claude Opus 4 8
Rank Model Wins Losses Ties Comparisons
1 Qwen 3.6 35B-A3B Q4_K_M 💡 71 53 4 128
2 Qwen 3.6 27B Q4_K_M 💡 67 43 9 119
3 Gemma 4 31B Instruct QAT 💡 51 63 7 121
4 Gemma 4 26B-A4B Instruct QAT 💡 41 79 9 129
5 GLM 5.2 💡 22 10 2 34
6 MiMo V2.5 Pro 💡 22 12 0 34
7 Gemma 4 26B-A4B Instruct Q8_0 💡 22 36 3 61
8 DeepSeek V4 Pro 💡 20 12 3 35
9 Qwen 3.7 Max 💡 20 12 2 34
10 Gemma 4 31B Instruct Q8_0 💡 20 35 6 61
11 DeepSeek V4 Flash 💡 16 19 0 35
12 MiMo V2.5 💡 15 17 0 32
13 Qwen 3.7 Plus 💡 14 16 3 33
14 MiniMax M3 💡 10 4 0 14

GPT-5.6 Sol

Partial — not a final ranking

651 validated duel verdicts

Results from GPT-5.6 Sol
Rank Model Wins Losses Ties Comparisons
1 Qwen 3.6 27B Q4_K_M 💡 69 53 0 122
2 Gemma 4 26B-A4B Instruct QAT 💡 63 63 2 128
3 GLM 5.2 💡 62 32 1 95
4 Qwen 3.6 35B-A3B Q4_K_M 💡 59 68 1 128
5 DeepSeek V4 Pro 💡 58 34 3 95
6 Gemma 4 31B Instruct QAT 💡 58 65 1 124
7 Qwen 3.7 Plus 💡 56 34 5 95
8 Qwen 3.7 Max 💡 52 39 4 95
9 DeepSeek V4 Flash 💡 34 60 1 95
10 MiMo V2.5 💡 31 64 0 95
11 Gemma 4 26B-A4B Instruct Q8_0 💡 30 31 0 61
12 Gemma 4 31B Instruct Q8_0 💡 30 31 0 61
13 MiMo V2.5 Pro 💡 28 63 4 95
14 MiniMax M3 💡 10 3 0 13