AI judge results

Each judge sees two anonymous answers to the same native English writing prompt. It picks the stronger text or declares a tie, without seeing the model names.

A judge may also evaluate a pair containing a response from its own model; the answer identities remain hidden while it decides.

The tables add up every validated, published duel for each judge: wins, losses, ties and total comparisons. Blinding keeps each verdict focused on the writing rather than the brand behind it.

Ranks are ordered first by total wins, not by win rate or opponent strength; when coverage differs, read wins, losses and ties together with comparisons.

For the selected writing categories, ranks are assigned before publisher, size, model-status, generation-mode and access filters hide models, so gaps can remain in the visible numbering.

AI judges make it possible to compare many model pairs with a consistent protocol. They do not replace readers: human votes are counted separately in the human ranking, revealing where people and machines agree or disagree.

Latest included verdict:

Claude Fable 5

1084 validated duel verdicts

Results from Claude Fable 5
Rank Model Wins Losses Ties Comparisons
1 Claude Fable 5 adaptive high 💡 758 68 79 905
2 MiniMax M3 enabled 💡 44 39 8 91
3 Claude Opus 5 adaptive high 💡 24 10 11 45
4 Qwen3.7 Plus enabled 💡 15 53 4 72
7 MiMo V2.5 Pro enabled 💡 14 53 5 72
8 DeepSeek V4 Pro enabled 💡 13 51 7 71
9 Kimi K3 enabled 💡 9 16 20 45
11 DeepSeek V4 Flash enabled 💡 9 58 5 72
13 MiMo V2.5 enabled 💡 7 62 3 72
14 Gemma 4 31B QAT-Q4_0 💡 🏠 6 29 1 36
15 Gemma 4 12B QAT-Q4_0 💡 🏠 4 23 1 28
17 Gemma 4 12B Q8_0 (non-reasoning) 🏠 3 10 4 17
18 Qwen3.6 35B A3B Q8_0 💡 🏠 3 31 3 37
19 GPT-5.6 Luna high 💡 3 39 3 45
20 GPT-5.6 Terra high 💡 3 39 3 45
22 Gemma 4 12B Q8_0 💡 🏠 2 30 6 38
23 GPT-5.6 Sol high 💡 2 36 6 44
24 Gemma 4 31B QAT-Q4_0 (non-reasoning) 🏠 1 4 1 6
25 Qwen3.6 35B A3B Q8_0 (non-reasoning) 🏠 1 11 5 17
26 Gemma 4 31B Q8_0 💡 🏠 1 28 4 33
27 Qwen3.6 35B A3B Q4_K_M 💡 🏠 1 30 3 34
28 Mistral Medium 3.5 128B Q4_K_M (non-reasoning) 🏠 0 8 0 8
29 Gemma 4 26B A4B Q8_0 💡 🏠 0 22 2 24
30 Gemma 4 26B A4B QAT-Q4_0 💡 🏠 0 31 0 31

Claude Opus 4.8

843 validated duel verdicts

Results from Claude Opus 4.8
Rank Model Wins Losses Ties Comparisons
2 Qwen3.6 35B A3B Q4_K_M 💡 🏠 78 68 5 151
3 MiniMax M3 enabled 💡 75 7 0 82
4 Gemma 4 31B QAT-Q4_0 💡 🏠 62 76 9 147
5 Gemma 4 26B A4B QAT-Q4_0 💡 🏠 56 106 9 171
6 Gemma 4 31B Q8_0 💡 🏠 38 93 7 138
9 MiMo V2.5 Pro enabled 💡 36 20 0 56
11 DeepSeek V4 Pro enabled 💡 35 18 3 56
12 Claude Fable 5 adaptive high 💡 33 0 0 33
13 DeepSeek V4 Flash enabled 💡 31 27 0 58
14 Qwen3.6 35B A3B Q8_0 💡 🏠 29 34 2 65
15 MiMo V2.5 enabled 💡 28 27 0 55
16 Qwen3.7 Plus enabled 💡 27 26 4 57
17 Gemma 4 26B A4B Q8_0 💡 🏠 27 48 3 78
18 Gemma 4 12B Q8_0 💡 🏠 21 42 1 64
19 Gemma 4 12B QAT-Q4_0 💡 🏠 19 66 3 88
21 GPT-5.6 Sol high 💡 5 1 1 7
22 Mistral Medium 3.5 128B Q4_K_M (non-reasoning) 🏠 5 6 0 11
23 Gemma 4 12B Q8_0 (non-reasoning) 🏠 5 9 0 14
24 Qwen3.6 35B A3B Q8_0 (non-reasoning) 🏠 5 10 0 15
25 Gemma 4 31B QAT-Q4_0 (non-reasoning) 🏠 0 3 0 3

Claude Opus 5

17197 validated duel verdicts

Results from Claude Opus 5
Rank Model Wins Losses Ties Comparisons
1 Claude Fable 5 adaptive high 💡 733 49 23 805
2 Claude Opus 5 adaptive high 💡 506 23 11 540
3 Kimi K3 enabled 💡 492 37 11 540
4 Claude Sonnet 5 adaptive high 💡 448 61 31 540
5 MiniMax M3 enabled 💡 425 98 19 542
6 Muse Glimmer 30B K-Quant Dynamic 💡 🏠 425 107 44 576
7 GPT-5.6 Sol high 💡 399 121 25 545
10 DeepSeek V4 Pro enabled 💡 364 165 16 545
11 Muse Glimmer 30B K-Quant 17GB 💡 🏠 360 177 37 574
12 DeepSeek V4 Flash enabled 💡 339 189 16 544
13 GPT-5.6 Terra high 💡 324 184 32 540
14 Gemma 4 31B QAT-Q4_0 💡 🏠 295 245 38 578
16 MiMo V2.5 Pro enabled 💡 286 226 30 542
17 Qwen3.6 35B A3B Q8_0 💡 🏠 285 266 41 592
18 GPT-5.6 Luna high 💡 278 233 29 540
19 MiMo V2.5 enabled 💡 278 242 22 542
21 Gemma 4 31B Q8_0 💡 🏠 256 257 56 569
22 Gemma 4 31B QAT-Q4_0 (non-reasoning) 🏠 248 306 48 602
23 Qwen3.7 Plus enabled 💡 234 280 28 542
24 Qwen3.6 35B A3B Q4_K_M 💡 🏠 233 315 30 578
27 Gemma 4 26B A4B QAT-Q4_0 💡 🏠 224 290 48 562
28 GLM 5.3 max 💡 219 8 12 239
29 GLM 5.3 Flash high 💡 216 10 12 238
30 Gemma 4 12B Q8_0 💡 🏠 216 315 39 570
31 Gemma 4 26B A4B Q8_0 💡 🏠 214 318 37 569
33 Gemma 4 12B Q8_0 (non-reasoning) 🏠 202 322 33 557
34 Gemma 4 12B QAT-Q4_0 💡 🏠 202 325 51 578
35 Qwen3.8 27B Q4_K_M 💡 🏠 188 68 22 278
36 Nemotron 3 Super 120B A12B Q4_K_M 💡 🏠 179 333 27 539
37 Hy3 high 💡 173 49 17 239
38 Mistral Medium 3.5 128B Q4_K_M (non-reasoning) 🏠 173 347 23 543
39 Qwen3.8 27B Q8_0 💡 🏠 172 69 28 269
40 Qwen3.8 Max xhigh 💡 168 49 21 238
41 Seed 2.1 Turbo (non-reasoning) disabled 160 72 9 241
43 Inkling high 💡 159 69 9 237
44 Muse Spark 1.2 high 💡 156 49 33 238
45 Aion 3.0 mandatory 💡 156 67 16 239
46 Qwen3.8 27B Q6_K 💡 🏠 155 92 36 283
47 Gemma 4 31B Q6_K 💡 🏠 147 97 30 274
48 Gemma 4 31B Q8_0 (non-reasoning) 🏠 146 100 43 289
49 Qwen3.6 35B A3B Q8_0 (non-reasoning) 🏠 146 398 11 555
50 Grok 4.6 high 💡 145 83 11 239
51 Gemma 4 26B A4B Q6_K 💡 🏠 143 96 38 277
52 Qwen3.8 Flash high 💡 142 88 8 238
53 Gemma 4 12B Q6_K 💡 🏠 142 112 30 284
54 Hy4 Preview high 💡 138 72 28 238
55 Gemini 3.7 Flash medium 💡 132 51 56 239
56 Gemma 4 31B Q6_K (non-reasoning) 🏠 125 117 35 277
57 Gemma 4 26B A4B Q6_K (non-reasoning) 🏠 124 116 34 274
58 Gemma 4 26B A4B QAT-Q4_0 (non-reasoning) 🏠 113 168 16 297
59 Gemma 4 E2B Q8_0 💡 🏠 106 423 28 557
60 Gemma 4 12B QAT-Q4_0 (non-reasoning) 🏠 95 171 14 280
61 Gemma 4 26B A4B Q8_0 (non-reasoning) 🏠 95 175 17 287
62 Qwen3.8 27B Q6_K (non-reasoning) 🏠 92 173 13 278
63 Qwen3.6 35B A3B Q4_K_M (non-reasoning) 🏠 91 185 6 282
64 Granite 4.2 8B Q6_K 💡 🏠 88 128 22 238
65 Nemotron 3 Nano Omni 30B A3B Q4_K_M 💡 🏠 88 448 5 541
66 Gemma 4 E4B Q8_0 💡 🏠 86 458 9 553
67 Granite 4.2 30B Q8_0 💡 🏠 83 131 23 237
68 Granite 4.2 30B Q6_K 💡 🏠 79 140 18 237
69 Gemma 4 12B Q6_K (non-reasoning) 🏠 75 190 11 276
70 Gemma 4 E4B Q8_0 (non-reasoning) 🏠 74 465 14 553
71 Mistral Medium 3.5 high 💡 71 155 13 239
72 Granite 4.2 8B Q8_0 💡 🏠 70 182 12 264
73 Qwen3.8 27B Q4_K_M (non-reasoning) 🏠 69 194 16 279
74 Gemma 4 E4B Q4_K_M (non-reasoning) 🏠 66 216 3 285
75 Qwen3.8 27B Q8_0 (non-reasoning) 🏠 60 216 7 283
76 Gemma 4 E4B Q6_K (non-reasoning) 🏠 59 215 5 279
77 Gemma 4 E4B Q6_K 💡 🏠 55 227 4 286
78 Gemma 4 E2B Q4_K_M 💡 🏠 55 491 12 558
79 Nemotron 3 Nano Omni 30B A3B Q8_0 💡 🏠 51 195 9 255
80 Granite 4.2 8B Q4_K_M 💡 🏠 48 174 16 238
81 Granite 4.2 30B Q4_K_M 💡 🏠 47 174 16 237
82 Gemma 4 E4B Q4_K_M 💡 🏠 43 219 6 268
83 Gemma 4 E2B Q8_0 (non-reasoning) 🏠 42 236 2 280
84 Gemma 4 E2B Q6_K 💡 🏠 41 226 17 284
85 Gemma 4 E2B Q4_K_M (non-reasoning) 🏠 24 260 4 288
86 Gemma 4 E2B Q6_K (non-reasoning) 🏠 21 259 1 281
87 Gemini 3.8 Flash High 💡 7 2 7 16

GPT-5.6 Sol

20474 validated duel verdicts

Results from GPT-5.6 Sol
Rank Model Wins Losses Ties Comparisons
1 Claude Fable 5 adaptive high 💡 1212 163 3 1378
2 GPT-5.6 Sol high 💡 802 103 9 914
3 Kimi K3 enabled 💡 488 63 2 553
4 DeepSeek V4 Flash enabled 💡 463 443 3 909
6 Claude Opus 5 adaptive high 💡 449 103 0 552
7 GPT-5.6 Terra high 💡 423 126 3 552
8 Claude Sonnet 5 adaptive high 💡 416 140 2 558
9 MiniMax M3 enabled 💡 401 193 0 594
10 GPT-5.6 Luna high 💡 400 148 4 552
11 DeepSeek V4 Pro enabled 💡 383 356 5 744
12 Qwen3.7 Plus enabled 💡 364 248 7 619
13 GLM 5.3 Flash high 💡 355 160 1 516
14 Muse Glimmer 30B K-Quant Dynamic 💡 🏠 352 177 6 535
15 Gemma 4 31B Q8_0 💡 🏠 351 260 8 619
16 Hy4 Preview high 💡 349 147 6 502
17 Qwen3.8 Flash high 💡 314 185 7 506
19 Gemma 4 31B QAT-Q4_0 💡 🏠 303 334 3 640
20 Gemma 4 12B QAT-Q4_0 💡 🏠 297 313 5 615
21 Gemma 4 26B A4B Q8_0 💡 🏠 291 345 3 639
22 MiMo V2.5 Pro enabled 💡 290 326 7 623
23 Qwen3.6 35B A3B Q8_0 💡 🏠 287 285 3 575
24 Qwen3.6 35B A3B Q4_K_M 💡 🏠 287 375 3 665
25 Gemma 4 26B A4B QAT-Q4_0 💡 🏠 287 386 4 677
26 Muse Glimmer 30B K-Quant 17GB 💡 🏠 281 255 2 538
27 GLM 5.3 max 💡 280 45 1 326
28 Muse Spark 1.2 high 💡 279 58 5 342
30 Gemma 4 31B QAT-Q4_0 (non-reasoning) 🏠 278 271 9 558
31 MiMo V2.5 enabled 💡 271 351 0 622
32 Hy3 high 💡 267 73 1 341
33 Gemma 4 12B Q8_0 💡 🏠 267 344 0 611
35 Gemini 3.7 Flash medium 💡 252 83 8 343
37 Qwen3.8 Max xhigh 💡 248 71 8 327
38 Gemma 4 12B Q8_0 (non-reasoning) 🏠 245 399 3 647
39 Aion 3.0 mandatory 💡 240 100 1 341
40 Grok 4.6 high 💡 232 93 1 326
41 Granite 4.2 30B Q8_0 💡 🏠 228 250 2 480
43 Nemotron 3 Super 120B A12B Q4_K_M 💡 🏠 205 350 4 559
45 Qwen3.8 27B Q8_0 💡 🏠 197 129 2 328
46 Gemma 4 26B A4B Q6_K 💡 🏠 195 123 1 319
47 Qwen3.8 27B Q4_K_M 💡 🏠 188 126 5 319
48 Qwen3.8 27B Q6_K 💡 🏠 187 132 0 319
49 Gemma 4 31B Q6_K (non-reasoning) 🏠 185 134 1 320
50 Gemma 4 31B Q6_K 💡 🏠 182 137 3 322
51 Gemma 4 26B A4B Q6_K (non-reasoning) 🏠 173 146 3 322
52 Gemma 4 31B Q8_0 (non-reasoning) 🏠 170 140 2 312
53 Gemma 4 E2B Q8_0 💡 🏠 167 378 1 546
54 Gemma 4 12B Q6_K 💡 🏠 163 159 0 322
55 Mistral Medium 3.5 128B Q4_K_M (non-reasoning) 🏠 156 650 4 810
56 Inkling high 💡 152 189 0 341
57 Granite 4.2 8B Q8_0 💡 🏠 146 319 3 468
58 Qwen3.6 35B A3B Q8_0 (non-reasoning) 🏠 146 484 0 630
59 Nemotron 3 Nano Omni 30B A3B Q8_0 💡 🏠 142 201 2 345
60 Gemma 4 26B A4B QAT-Q4_0 (non-reasoning) 🏠 140 157 3 300
61 Mistral Medium 3.5 high 💡 134 197 1 332
62 Granite 4.2 8B Q6_K 💡 🏠 134 346 0 480
63 Granite 4.2 8B Q4_K_M 💡 🏠 134 348 1 483
64 Granite 4.2 30B Q6_K 💡 🏠 133 185 0 318
65 Gemma 4 E2B Q4_K_M 💡 🏠 130 414 0 544
66 Seed 2.1 Turbo (non-reasoning) disabled 124 214 0 338
67 Gemma 4 12B QAT-Q4_0 (non-reasoning) 🏠 116 199 2 317
68 Gemma 4 E4B Q8_0 (non-reasoning) 🏠 113 423 0 536
70 Gemma 4 26B A4B Q8_0 (non-reasoning) 🏠 105 207 2 314
71 Granite 4.2 30B Q4_K_M 💡 🏠 105 370 1 476
72 Gemma 4 12B Q6_K (non-reasoning) 🏠 103 225 1 329
73 Gemma 4 E4B Q8_0 💡 🏠 103 432 1 536
74 Nemotron 3 Nano Omni 30B A3B Q4_K_M 💡 🏠 101 456 0 557
75 Qwen3.8 27B Q8_0 (non-reasoning) 🏠 98 218 0 316
76 Gemma 4 E4B Q6_K 💡 🏠 97 214 0 311
77 Qwen3.8 27B Q6_K (non-reasoning) 🏠 86 234 0 320
78 Gemma 4 E4B Q4_K_M (non-reasoning) 🏠 79 234 0 313
79 Qwen3.6 35B A3B Q4_K_M (non-reasoning) 🏠 75 242 0 317
80 Gemma 4 E2B Q6_K 💡 🏠 73 246 0 319
81 Gemma 4 E4B Q6_K (non-reasoning) 🏠 71 247 0 318
82 Qwen3.8 27B Q4_K_M (non-reasoning) 🏠 71 248 0 319
83 Gemma 4 E4B Q4_K_M 💡 🏠 71 258 0 329
84 Gemma 4 E2B Q4_K_M (non-reasoning) 🏠 65 247 0 312
85 Gemma 4 E2B Q8_0 (non-reasoning) 🏠 64 255 0 319
86 Gemini 3.8 Flash High 💡 43 2 2 47
87 Gemma 4 E2B Q6_K (non-reasoning) 🏠 37 281 0 318
88 Claude Fable 5.1 High 💡 11 7 0 18

The AIs have had their say. Your turn.

An AI judge can read thousands of duels without coffee or a lunch break. It still cannot tell us what makes you laugh, moves you, or keeps you reading. Your blind vote feeds a separate human ranking and helps reveal where human taste and machine judgment part ways.

Judge two anonymous texts