About this project
About the author
My name is Johannes Eva, and I have been building websites for twenty-seven years. The AI Writing Benchmark began with a gap: no existing leaderboard measured what interested me, the writing quality of AI models in English and other languages.
Public benchmarks mainly focus on code, mathematics and reasoning. The small part devoted to writing is usually evaluated only in English. Here, only the quality of the generated texts counts. The prompts are written directly in each language, and the answers are compared blindly.
Models that can run at home have a special place here. Gemma 4, Qwen 3.6 and Mistral Medium 3.5, among others, are marked with 🏠. I ran them locally in several quantisations.
This site is a personal project. No model provider pays me, and I have no reason to favour one model over another. The published results come from anonymous duels.
Hardware
The computers behind the local models
The main machine is a Framework Desktop built around AMD’s high-memory Strix Halo platform. Its unified memory makes it possible to test local models that do not fit on a conventional consumer graphics card.
- Computer
- Framework Desktop, AMD Ryzen AI Max 300 Series, revision A6
- Processor
- AMD Ryzen AI Max+ 395, 16 cores and 32 threads, up to 5.19 GHz
- Graphics
- AMD Radeon 8060S integrated graphics
- Memory
- 128 GB of unified memory, shared by the processor and graphics unit
- Storage
- Samsung 9100 PRO 2 TB and WD_BLACK SN770 2 TB NVMe SSDs
- Second machine
- A separate desktop with an NVIDIA GeForce RTX 5070 Ti