About the author

My name is Johannes Eva, and I have been building websites for 27 years. The AI Writing Benchmark began with a gap: no existing leaderboard measured what interested me, the writing quality of AI models in English and other languages.

Public benchmarks mainly focus on code, mathematics and reasoning. The small part devoted to writing is usually evaluated only in English. Here, only the quality of the generated texts counts. The prompts are written directly in each language, and the answers are compared blindly.

Models that can run at home have a special place here. Gemma 4, Qwen 3.6 and Mistral Medium 3.5, among others, are marked with 🏠. I ran them locally in several quantisations.

This site is a personal project. No model provider pays me, and I have no reason to favour one model over another.

Framework Desktop used to run local language models for the AI Writing Benchmark

The machine used for local models

Local models run on a Framework Desktop with an AMD Ryzen AI Max+ 395. Its 128 GB of unified memory is available to both the processor and graphics unit, making it possible to load models larger than most consumer graphics cards can hold.

Computer
Framework Desktop, AMD Ryzen AI Max 300 Series, revision A6
Processor
AMD Ryzen AI Max+ 395, 16 cores and 32 threads, up to 5.19 GHz
Graphics
AMD Radeon 8060S integrated graphics
Memory
128 GB of unified memory, shared by the processor and graphics unit

Observed local generation speed

Median tokens generated per second on this machine across valid benchmark responses, including reasoning tokens. Speed varies with context length and software.

Observed local generation speed
Model Quantisation Median speed Responses
Qwen3.6 35B A3B Q4_K_M 71.2 tokens/s 379
Gemma 4 26B A4B QAT QAT-Q4_0 61.3 tokens/s 488
Qwen3.6 35B A3B Q8_0 50.1 tokens/s 425
Gemma 4 26B A4B Q8_0 40.4 tokens/s 382
Gemma 4 E2B Q4_K_M 39.7 tokens/s 189
Gemma 4 E4B Q8_0 35.9 tokens/s 542
Gemma 4 E2B Q6_K 33.8 tokens/s 108
Gemma 4 E2B Q8_0 29.5 tokens/s 152
Nemotron 3 Nano Omni Q4_K_M 25.1 tokens/s 960
Gemma 4 E4B Q4_K_M 22.9 tokens/s 111
Nemotron 3 Super Q4_K_M 21.8 tokens/s 226
Gemma 4 26B A4B Q6_K 21.4 tokens/s 116
Laguna S 2.1 UD Q4_K_M 21.4 tokens/s 371
Nemotron 3 Nano Omni Q8_0 21.1 tokens/s 72
Gemma 4 E4B Q6_K 18.9 tokens/s 108
Gemma 4 12B Q8_0 14.5 tokens/s 429
Gemma 4 12B QAT QAT-Q4_0 11.2 tokens/s 1048
Gemma 4 31B QAT QAT-Q4_0 10.6 tokens/s 499
Gemma 4 12B Q6_K 9.0 tokens/s 144
Qwen3.8 27B Q4_K_M 7.6 tokens/s 96
Qwen3.8 27B Q6_K 6.5 tokens/s 72
Gemma 4 31B Q8_0 6.2 tokens/s 299
Muse Glimmer 30B Kquant 17gb K-Quant 17GB 5.5 tokens/s 138
Muse Glimmer 30B Kquant Dynamic K-Quant Dynamic 4.8 tokens/s 138
Qwen3.8 27B Q8_0 3.8 tokens/s 75
Gemma 4 31B Q6_K 3.5 tokens/s 108
Mistral Medium 3.5 128B Q4_K_M 2.8 tokens/s 771

Comparable published benchmarks for Ryzen AI Max+ 395