Our updated writing test compares both Muse Glimmer quantizations with Gemma, Qwen and Nemotron across four languages, then weighs memory, speed and factual risk.
The models that write best in English
The AI Writing Benchmark measures how well leading AI models write in English. Responses to a shared corpus of English writing prompts are compared blind, without model names, and the verdicts are combined into a ranking for English.
Leading across all categories
The 10 best models
| # | Model | Score |
|---|---|---|
| 1 | 94.1 | |
| 2 | 93.7 | |
| 3 | 93.5 | |
| 4 | 92.8 | |
| 5 | 90.4 | |
| 6 | 88.4 | |
| 7 | 87.0 | |
| 8 | 82.5 | |
| 9 | 82.4 | |
| 10 | 81.4 |
This score is not a percentage grade: 50 means an estimated one-in-two chance of beating the average model in English.
Last updated:
See the full leaderboardThe 10 best local variants
Estimated VRAM is the approximate graphics memory needed to load the entire model.
| # | Model | Score | Estimated VRAM |
|---|---|---|---|
| 1 | 79.3 | β 32 GB | |
| 2 | 67.2 | β 48 GB | |
| 3 | 67.1 | β 24 GB | |
| 4 | 65.2 | β 24 GB | |
| 5 | 64.9 | β 32 GB | |
| 6 | 61.0 | β 32 GB | |
| 7 | 59.4 | β 32 GB | |
| 8 | 58.5 | β 48 GB | |
| 9 | 58.5 | β 24 GB | |
| 10 | 58.4 | β 32 GB |
Some software can keep part of the model in system memory. This reduces the VRAM required, but usually makes generation slower.
The ranking evaluates writing quality alone. Prompts are written directly in the language being tested, and model names remain hidden until the verdict so reputation cannot influence the evaluation. This site is a personal project. No model provider pays me, and I have no reason to favour one model over another.
Explore the results
In the current direct sample, reasoning-enabled responses earned 66.2% of the available points in English. The uncertainty interval still includes parity, the result differs sharply from one model to the next, and no human votes have compared the two settings directly yet.