The models that write best in English

The AI Writing Benchmark measures how well leading AI models write in English. Responses to a shared corpus of English writing prompts are compared blind, without model names, and the verdicts are combined into a ranking for English.

Leading across all categories

The 10 best models

The ten best models for writing in English
#ModelScore
1Kimi K3 πŸ’‘94.1
2Claude Fable 5 High πŸ’‘93.7
3Claude Opus 5 πŸ’‘93.5
4GLM 5.3 πŸ’‘92.8
5Claude Opus 4.8 πŸ’‘90.4
6GPT-5.6 Sol πŸ’‘88.4
7Claude Sonnet 5 πŸ’‘87.0
8Muse Spark 1.2 πŸ’‘82.5
9MiniMax M3 πŸ’‘82.4
10Hy3 πŸ’‘81.4

This score is not a percentage grade: 50 means an estimated one-in-two chance of beating the average model in English.

Last updated:

See the full leaderboard

The 10 best local variants

Estimated VRAM is the approximate graphics memory needed to load the entire model.

The ten best exact variants run locally
#ModelScoreEstimated VRAM
1Muse Glimmer 30B K-Quant Dynamic πŸ’‘ 🏠79.3β‰ˆ 32 GB
2Qwen3.8 27B Q8_0 πŸ’‘ 🏠67.2β‰ˆ 48 GB
3Qwen3.8 27B Q4_K_M πŸ’‘ 🏠67.1β‰ˆ 24 GB
4Muse Glimmer 30B K-Quant 17GB πŸ’‘ 🏠65.2β‰ˆ 24 GB
5Qwen3.6 27B Q8_0 πŸ’‘ 🏠64.9β‰ˆ 32 GB
6Gemma 4 26B A4B Q6_K πŸ’‘ 🏠61.0β‰ˆ 32 GB
7Qwen3.8 27B Q6_K πŸ’‘ 🏠59.4β‰ˆ 32 GB
8Gemma 4 31B Q8_0 πŸ’‘ 🏠58.5β‰ˆ 48 GB
9Qwen3.6 27B Q4_K_M πŸ’‘ 🏠58.5β‰ˆ 24 GB
10Gemma 4 31B Q6_K πŸ’‘ 🏠58.4β‰ˆ 32 GB

Some software can keep part of the model in system memory. This reduces the VRAM required, but usually makes generation slower.

86models compared
38746published AI verdicts

The ranking evaluates writing quality alone. Prompts are written directly in the language being tested, and model names remain hidden until the verdict so reputation cannot influence the evaluation. This site is a personal project. No model provider pays me, and I have no reason to favour one model over another.

Explore the results

Does AI Write Better in English With or Without Reasoning?

In the current direct sample, reasoning-enabled responses earned 66.2% of the available points in English. The uncertainty interval still includes parity, the result differs sharply from one model to the next, and no human votes have compared the two settings directly yet.