AI judge results

Each judge sees two anonymous answers to the same native English writing prompt. It picks the stronger text—or a tie—without seeing the model names.

The tables add up every validated, published duel for each judge: wins, losses, ties and total comparisons. Blinding keeps each verdict focused on the writing rather than the brand behind it.

AI judges let us compare many model pairs with a consistent protocol. They do not replace readers: human votes are counted separately in the human ranking, so you can see where people and machines agree—or do not.

Latest included verdict:

Claude Fable 5

225 validated duel verdicts

Results from Claude Fable 5
Rank Model Wins Losses Ties Comparisons
1 Claude Fable 5 High 💡 45 0 1 46
2 MiniMax M3 💡 42 5 2 49
3 Qwen 3.7 Plus 💡 15 11 4 30
4 Qwen 3.7 Max 💡 13 11 6 30
5 MiMo V2.5 Pro 💡 13 14 4 31
6 DeepSeek V4 Pro 💡 12 12 6 30
7 Qwen 3.6 27B Q8_0 💡 9 7 2 18
8 DeepSeek V4 Flash 💡 9 19 3 31
9 MiMo V2.5 💡 7 21 3 31
10 GLM 5.2 💡 6 20 4 30
11 Gemma 4 31B Instruct QAT 💡 5 8 1 14
12 Gemma 4 12B Instruct Q8_0 (non-reasoning) 3 2 4 9
13 Qwen 3.6 27B Q4_K_M 💡 3 8 2 13
14 Qwen 3.6 27B Q8_0 (non-reasoning) 2 4 4 10
15 Gemma 4 12B Instruct Q8_0 💡 2 8 6 16
16 Qwen 3.6 35B-A3B Q8_0 💡 2 10 3 15
17 Qwen 3.6 35B-A3B Q8_0 (non-reasoning) 1 3 5 9
18 Gemma 4 12B Instruct QAT 💡 1 3 0 4
19 Gemma 4 31B Instruct Q8_0 💡 1 5 3 9
20 Qwen 3.6 35B-A3B Q4_K_M 💡 1 7 3 11
21 Gemma 4 26B-A4B Instruct Q8_0 💡 0 4 0 4
22 Gemma 4 26B-A4B Instruct QAT 💡 0 10 0 10

Claude Opus 4 8

559 validated duel verdicts

Results from Claude Opus 4 8
Rank Model Wins Losses Ties Comparisons
1 Qwen 3.6 27B Q4_K_M 💡 71 47 9 127
2 Qwen 3.6 35B-A3B Q4_K_M 💡 71 58 5 134
3 Gemma 4 31B Instruct QAT 💡 53 69 7 129
4 MiniMax M3 💡 43 4 0 47
5 Gemma 4 26B-A4B Instruct QAT 💡 42 84 9 135
6 MiMo V2.5 Pro 💡 27 16 0 43
7 DeepSeek V4 Pro 💡 26 14 3 43
8 GLM 5.2 💡 26 15 2 43
9 Qwen 3.7 Max 💡 23 17 2 42
10 Gemma 4 26B-A4B Instruct Q8_0 💡 22 38 3 63
11 DeepSeek V4 Flash 💡 20 23 0 43
12 Gemma 4 31B Instruct Q8_0 💡 20 41 6 67
13 Qwen 3.7 Plus 💡 18 20 3 41
14 MiMo V2.5 💡 18 23 0 41
15 Qwen 3.6 27B Q8_0 💡 15 13 0 28
16 Qwen 3.6 35B-A3B Q8_0 💡 11 15 1 27
17 Gemma 4 12B Instruct Q8_0 💡 9 18 0 27
18 Gemma 4 12B Instruct QAT 💡 6 4 0 10
19 Qwen 3.6 35B-A3B Q8_0 (non-reasoning) 5 4 0 9
20 Qwen 3.6 27B Q8_0 (non-reasoning) 5 5 0 10
21 Gemma 4 12B Instruct Q8_0 (non-reasoning) 3 6 0 9

GPT-5.6 Sol

1296 validated duel verdicts

Results from GPT-5.6 Sol
Rank Model Wins Losses Ties Comparisons
1 GLM 5.2 💡 108 47 1 156
2 DeepSeek V4 Pro 💡 97 55 3 155
3 Qwen 3.7 Plus 💡 91 59 5 155
4 Qwen 3.7 Max 💡 91 60 4 155
5 Qwen 3.6 27B Q4_K_M 💡 89 65 0 154
6 Gemma 4 26B-A4B Instruct QAT 💡 83 79 3 165
7 Qwen 3.6 35B-A3B Q4_K_M 💡 80 85 1 166
8 Gemma 4 31B Instruct QAT 💡 78 76 2 156
9 MiMo V2.5 Pro 💡 68 83 4 155
10 MiniMax M3 💡 66 29 0 95
11 DeepSeek V4 Flash 💡 63 91 1 155
12 MiMo V2.5 💡 54 101 0 155
13 Gemma 4 31B Instruct Q8_0 💡 51 42 0 93
14 Qwen 3.6 35B-A3B Q8_0 💡 47 41 0 88
15 Gemma 4 26B-A4B Instruct Q8_0 💡 42 46 1 89
16 Qwen 3.6 27B Q8_0 💡 42 46 0 88
17 Gemma 4 12B Instruct Q8_0 💡 37 51 0 88
18 Gemma 4 12B Instruct Q8_0 (non-reasoning) 37 59 2 98
19 Qwen 3.6 27B Q8_0 (non-reasoning) 21 77 0 98
20 Gpt 5.6 Sol 💡 19 3 0 22
21 Qwen 3.6 35B-A3B Q8_0 (non-reasoning) 16 82 0 98
22 Gemma 4 12B Instruct QAT 💡 2 5 1 8

The AIs have had their say. Your turn.

An AI judge can read thousands of duels without coffee or a lunch break. It still cannot tell us what makes you laugh, moves you, or keeps you reading. Your blind vote feeds a separate human ranking and helps reveal where human taste and machine judgment part ways.

Judge two anonymous texts