About the author
My name is Johannes Eva, and I have been building websites for 27 years. The AI Writing Benchmark began with a gap: no existing leaderboard measured what interested me, the writing quality of AI models in English and other languages.
Public benchmarks mainly focus on code, mathematics and reasoning. The small part devoted to writing is usually evaluated only in English. Here, only the quality of the generated texts counts. The prompts are written directly in each language, and the answers are compared blindly.
Models that can run at home have a special place here. Gemma 4, Qwen 3.6 and Mistral Medium 3.5, among others, are marked with 🏠. I ran them locally in several quantisations.
This site is a personal project. No model provider pays me, and I have no reason to favour one model over another.
The machine used for local models
Local models run on a Framework Desktop with an AMD Ryzen AI Max+ 395. Its 128 GB of unified memory is available to both the processor and graphics unit, making it possible to load models larger than most consumer graphics cards can hold.
- Computer
- Framework Desktop, AMD Ryzen AI Max 300 Series, revision A6
- Processor
- AMD Ryzen AI Max+ 395, 16 cores and 32 threads, up to 5.19 GHz
- Graphics
- AMD Radeon 8060S integrated graphics
- Memory
- 128 GB of unified memory, shared by the processor and graphics unit
Observed local generation speed
Median tokens generated per second on this machine across valid benchmark responses, including reasoning tokens. Speed varies with context length and software.
| Model | Quantisation | Median speed | Responses |
|---|---|---|---|
| Qwen3.6 35B A3B | Q4_K_M | 71.2 tokens/s | 379 |
| Gemma 4 26B A4B QAT | QAT-Q4_0 | 61.3 tokens/s | 488 |
| Qwen3.6 35B A3B | Q8_0 | 50.1 tokens/s | 425 |
| Gemma 4 26B A4B | Q8_0 | 40.4 tokens/s | 382 |
| Gemma 4 E2B | Q4_K_M | 39.7 tokens/s | 189 |
| Gemma 4 E4B | Q8_0 | 35.9 tokens/s | 542 |
| Gemma 4 E2B | Q6_K | 33.8 tokens/s | 108 |
| Gemma 4 E2B | Q8_0 | 29.5 tokens/s | 152 |
| Nemotron 3 Nano Omni | Q4_K_M | 25.1 tokens/s | 960 |
| Gemma 4 E4B | Q4_K_M | 22.9 tokens/s | 111 |
| Nemotron 3 Super | Q4_K_M | 21.8 tokens/s | 226 |
| Gemma 4 26B A4B | Q6_K | 21.4 tokens/s | 116 |
| Laguna S 2.1 UD | Q4_K_M | 21.4 tokens/s | 371 |
| Nemotron 3 Nano Omni | Q8_0 | 21.1 tokens/s | 72 |
| Gemma 4 E4B | Q6_K | 18.9 tokens/s | 108 |
| Gemma 4 12B | Q8_0 | 14.5 tokens/s | 429 |
| Gemma 4 12B QAT | QAT-Q4_0 | 11.2 tokens/s | 1048 |
| Gemma 4 31B QAT | QAT-Q4_0 | 10.6 tokens/s | 499 |
| Gemma 4 12B | Q6_K | 9.0 tokens/s | 144 |
| Qwen3.8 27B | Q4_K_M | 7.6 tokens/s | 96 |
| Qwen3.8 27B | Q6_K | 6.5 tokens/s | 72 |
| Gemma 4 31B | Q8_0 | 6.2 tokens/s | 299 |
| Muse Glimmer 30B Kquant 17gb | K-Quant 17GB | 5.5 tokens/s | 138 |
| Muse Glimmer 30B Kquant Dynamic | K-Quant Dynamic | 4.8 tokens/s | 138 |
| Qwen3.8 27B | Q8_0 | 3.8 tokens/s | 75 |
| Gemma 4 31B | Q6_K | 3.5 tokens/s | 108 |
| Mistral Medium 3.5 128B | Q4_K_M | 2.8 tokens/s | 771 |