Local AI: the best open-weight LLMs for writing in English

A permanent, continuously updated ranking of open-weight models under 256B parameters that you can actually run on your own machine — scored on English writing quality, not on coding or maths benchmarks.

Local AI ranking

Scores from the main English leaderboard. Ranks are calculated across all eligible local variants before display filters are applied; hidden variants can therefore leave gaps in the rank numbers.

Last updated:

Local model ranking by writing quality
RankModelAI judge scoreEstimated memory
1 Muse Glimmer 30B K-Quant Dynamic 💡 high 🏠 76.8 ≈ 32 GB
2 Qwen3.8 27B Q4_K_M 💡 xhigh** 🏠 66.9 ≈ 24 GB
3 Muse Glimmer 30B K-Quant 17GB 💡 high 🏠 64.1 ≈ 24 GB
5 Qwen3.8 27B Q8_0 💡 xhigh** 🏠 60.6 ≈ 48 GB
8 Gemma 4 31B Q8_0 💡 on* 🏠 56.3 ≈ 48 GB
11 Gemma 4 31B QAT-Q4_0 💡 on* 🏠 54.9 ≈ 24 GB
12 Qwen3.6 35B A3B Q8_0 💡 on* 🏠 54.2 ≈ 48 GB
13 Gemma 4 31B QAT-Q4_0 🏠 51.9 ≈ 24 GB
16 Gemma 4 12B QAT-Q4_0 💡 on* 🏠 46.8 ≈ 12 GB
17 Gemma 4 31B Q8_0 🏠 46.2 ≈ 48 GB
19 Gemma 4 26B A4B Q8_0 💡 on* 🏠 46.0 ≈ 32 GB
20 Qwen3.6 35B A3B Q4_K_M 💡 on* 🏠 45.3 ≈ 32 GB
21 Gemma 4 26B A4B QAT-Q4_0 💡 on* 🏠 44.3 ≈ 24 GB
22 Gemma 4 12B Q8_0 💡 on* 🏠 44.0 ≈ 16 GB
23 Laguna S 2.1 UD Q4_K_M 💡 on* 🏠 42.6 ≈ 80 GB
24 Gemma 4 26B A4B QAT-Q4_0 🏠 39.6 ≈ 24 GB
25 Gemma 4 12B Q8_0 🏠 39.5 ≈ 16 GB
26 Granite 4.2 30B Q8_0 💡 full 🏠 34.2 ≈ 48 GB
27 Nemotron 3 Super 120B A12B Q4_K_M 💡 full 🏠 31.8 ≈ 96 GB
28 Gemma 4 26B A4B Q8_0 🏠 30.0 ≈ 32 GB
31 Gemma 4 12B QAT-Q4_0 🏠 27.1 ≈ 12 GB
33 Granite 4.2 8B Q8_0 💡 full 🏠 26.4 ≈ 12 GB
34 Mistral Medium 3.5 💡 high 25.7
35 Mistral Medium 3.5 128B Q4_K_M 🏠 25.5 ≈ 80 GB
36 Granite 4.2 8B Q4_K_M 💡 full 🏠 24.6 ≈ 8 GB
37 Qwen3.6 35B A3B Q8_0 🏠 23.9 ≈ 48 GB
38 Qwen3.6 35B A3B Q4_K_M 🏠 23.4 ≈ 32 GB
39 Nemotron 3 Nano Omni 30B A3B Q8_0 💡 on* 🏠 23.1 ≈ 48 GB
41 Gemma 4 E2B Q8_0 💡 on* 🏠 21.0 ≈ 8 GB
42 Qwen3.8 27B Q8_0 🏠 21.0 ≈ 48 GB
43 Qwen3.8 27B Q4_K_M 🏠 20.4 ≈ 24 GB
44 Gemma 4 E4B Q4_K_M 🏠 19.3 ≈ 12 GB
47 Gemma 4 E4B Q8_0 🏠 15.0 ≈ 12 GB
48 Granite 4.2 30B Q4_K_M 💡 full 🏠 14.0 ≈ 24 GB
49 Gemma 4 E4B Q8_0 💡 on* 🏠 13.6 ≈ 12 GB
51 Nemotron 3 Nano Omni 30B A3B Q4_K_M 💡 on* 🏠 12.9 ≈ 32 GB
52 Gemma 4 E2B Q4_K_M 💡 on* 🏠 12.6 ≈ 8 GB
53 Gemma 4 E4B Q4_K_M 💡 on* 🏠 11.9 ≈ 12 GB
54 Gemma 4 E2B Q8_0 🏠 11.3 ≈ 8 GB
55 Gemma 4 E2B Q4_K_M 🏠 8.6 ≈ 8 GB

This score uses the main leaderboard scale, not a percentage grade: for the active category selection, 50 means an estimated one-in-two chance of beating the average English-writing model.

56 ranked local variants53995 published AI verdicts in the main leaderboard

Direct duel results

Each cell shows the share of available points earned by the row model against the column model: a win is worth 100%, a tie 50%, and a loss 0%. This result matrix uses only the compatible evaluation protocols included in the main ranking.

Directional matrix of points earned in direct duels. Values above 50% favor the row model; values below 50% favor the column model.

💡 marks a model with reasoning, ↯ one without reasoning. 🏠 marks a model run locally on the site author’s computer. “on*” means reasoning is enabled or disabled (local models with no intermediate levels): the requested fine-grained setting is not supported by the local file and falls back to the default enabled mode. For local models with adjustable reasoning (Qwen3.8, Muse Glimmer, Granite 4.2, Nemotron 3 Super), the displayed level (xhigh**, high, full) is the model’s default level: the requested fine-grained setting is not supported by the local file, which falls back to its default behavior.

View as a data table
Points earned in direct duels by model pair
ModelMuse Glimmer 30B K-Quant Dynamic 💡 high 🏠Qwen3.8 27B Q4_K_M 💡 xhigh** 🏠Muse Glimmer 30B K-Quant 17GB 💡 high 🏠Qwen3.8 27B Q8_0 💡 xhigh** 🏠Gemma 4 31B Q8_0 💡 on* 🏠Gemma 4 31B QAT-Q4_0 💡 on* 🏠Qwen3.6 35B A3B Q8_0 💡 on* 🏠Gemma 4 31B QAT-Q4_0 🏠Gemma 4 12B QAT-Q4_0 💡 on* 🏠Gemma 4 31B Q8_0 🏠Gemma 4 26B A4B Q8_0 💡 on* 🏠Qwen3.6 35B A3B Q4_K_M 💡 on* 🏠Gemma 4 26B A4B QAT-Q4_0 💡 on* 🏠Gemma 4 12B Q8_0 💡 on* 🏠Laguna S 2.1 UD Q4_K_M 💡 on* 🏠Gemma 4 26B A4B QAT-Q4_0 🏠Gemma 4 12B Q8_0 🏠Granite 4.2 30B Q8_0 💡 full 🏠Nemotron 3 Super 120B A12B Q4_K_M 💡 full 🏠Gemma 4 26B A4B Q8_0 🏠Gemma 4 12B QAT-Q4_0 🏠Granite 4.2 8B Q8_0 💡 full 🏠Mistral Medium 3.5 💡 highMistral Medium 3.5 128B Q4_K_M 🏠Granite 4.2 8B Q4_K_M 💡 full 🏠Qwen3.6 35B A3B Q8_0 🏠Qwen3.6 35B A3B Q4_K_M 🏠Nemotron 3 Nano Omni 30B A3B Q8_0 💡 on* 🏠Gemma 4 E2B Q8_0 💡 on* 🏠Qwen3.8 27B Q8_0 🏠Qwen3.8 27B Q4_K_M 🏠Gemma 4 E4B Q4_K_M 🏠Gemma 4 E4B Q8_0 🏠Granite 4.2 30B Q4_K_M 💡 full 🏠Gemma 4 E4B Q8_0 💡 on* 🏠Nemotron 3 Nano Omni 30B A3B Q4_K_M 💡 on* 🏠Gemma 4 E2B Q4_K_M 💡 on* 🏠Gemma 4 E4B Q4_K_M 💡 on* 🏠Gemma 4 E2B Q8_0 🏠Gemma 4 E2B Q4_K_M 🏠
Muse Glimmer 30B K-Quant Dynamic 💡 high 🏠53.8%63.5%83.3%77.5%72.5%80.0%87.5%80.0%90.9%87.5%81.0%83.3%85.0%81.0%100.0%90.0%75.0%78.6%81.8%100.0%84.6%100.0%82.5%80.8%95.0%90.9%91.7%90.5%100.0%63.6%83.3%85.0%100.0%100.0%100.0%95.2%83.3%90.9%90.9%
Qwen3.8 27B Q4_K_M 💡 xhigh** 🏠46.2%53.8%45.5%50.0%63.6%59.1%59.1%72.7%63.6%63.6%81.8%77.3%72.7%84.6%68.2%86.4%83.3%73.1%81.8%72.7%91.7%91.7%100.0%87.5%100.0%81.8%92.3%86.4%90.9%87.5%90.9%90.9%100.0%90.9%88.5%100.0%100.0%90.9%100.0%
Muse Glimmer 30B K-Quant 17GB 💡 high 🏠36.5%46.2%45.8%64.3%57.5%57.5%85.0%72.5%63.6%60.0%57.1%57.1%60.0%67.5%90.9%82.5%85.7%57.1%90.9%90.9%84.6%91.7%83.3%92.3%90.0%81.8%91.7%76.2%100.0%73.1%91.7%77.5%84.6%95.0%95.0%81.0%75.0%81.8%90.9%
Qwen3.8 27B Q8_0 💡 xhigh** 🏠16.7%54.5%54.2%45.5%63.6%59.1%54.5%54.2%63.6%72.7%77.3%84.6%66.7%45.8%77.8%77.3%75.0%75.0%80.0%90.9%81.2%100.0%70.8%78.1%63.6%77.3%72.7%83.3%75.0%72.7%72.7%68.2%90.6%100.0%91.7%100.0%100.0%90.9%100.0%
Gemma 4 31B Q8_0 💡 on* 🏠22.5%50.0%35.7%54.5%50.0%48.0%45.0%67.2%70.8%68.6%55.4%59.6%64.0%70.0%81.8%77.5%80.6%80.0%68.2%90.9%100.0%83.3%74.0%84.4%92.5%90.9%89.3%85.0%87.5%90.9%90.9%100.0%93.3%100.0%90.0%95.0%100.0%100.0%100.0%
Gemma 4 31B QAT-Q4_0 💡 on* 🏠27.5%36.4%42.5%36.4%50.0%50.0%67.5%73.1%68.2%58.3%54.8%54.0%65.2%62.5%68.2%65.8%67.6%75.0%86.4%77.3%92.3%86.4%77.5%76.7%76.3%100.0%72.7%77.5%83.3%90.9%85.0%92.1%86.7%78.9%90.0%87.5%100.0%100.0%90.9%
Qwen3.6 35B A3B Q8_0 💡 on* 🏠20.0%40.9%42.5%40.9%52.0%50.0%44.7%88.0%59.1%50.0%41.3%57.7%60.9%75.0%75.0%63.2%60.7%67.5%90.0%80.0%90.9%72.6%77.5%67.9%87.5%90.0%80.0%82.5%75.0%79.2%80.0%89.5%96.4%89.5%95.0%77.5%100.0%100.0%90.0%
Gemma 4 31B QAT-Q4_0 🏠12.5%40.9%15.0%45.5%55.0%32.5%55.3%62.5%63.6%65.0%60.5%55.3%65.8%60.0%45.5%47.4%69.2%85.0%63.6%86.4%69.2%77.3%80.0%73.1%89.5%81.8%72.7%72.5%90.9%68.2%86.4%89.5%92.3%100.0%95.0%85.0%90.9%100.0%100.0%
Gemma 4 12B QAT-Q4_0 💡 on* 🏠20.0%27.3%27.5%45.8%32.8%26.9%12.0%37.5%62.5%50.0%67.2%47.5%48.0%52.5%75.0%57.5%57.7%70.0%100.0%77.8%75.0%55.6%72.5%82.1%67.5%88.9%88.9%87.5%83.3%86.4%83.3%90.0%85.7%95.0%80.0%100.0%100.0%94.4%100.0%
Gemma 4 31B Q8_0 🏠9.1%36.4%36.4%36.4%29.2%31.8%40.9%36.4%37.5%31.8%45.8%45.8%72.7%40.9%40.9%90.9%73.1%81.8%81.8%72.7%61.5%75.0%72.7%69.2%86.4%81.8%84.6%77.3%81.8%81.8%90.9%77.3%83.3%81.8%90.9%100.0%95.5%100.0%100.0%
Gemma 4 26B A4B Q8_0 💡 on* 🏠12.5%36.4%40.0%27.3%31.4%41.7%50.0%35.0%50.0%68.2%53.8%72.2%50.0%57.5%75.0%57.5%69.2%75.0%88.9%55.6%76.7%77.8%66.7%63.3%87.5%95.0%83.3%80.0%72.7%77.3%83.3%95.0%92.9%100.0%90.0%95.0%100.0%83.3%90.0%
Qwen3.6 35B A3B Q4_K_M 💡 on* 🏠19.0%18.2%42.9%22.7%44.6%45.2%58.7%39.5%32.8%54.2%46.2%51.5%52.3%55.0%55.6%40.0%64.3%55.0%72.2%66.7%69.2%72.2%57.5%78.6%70.0%66.7%77.8%87.5%45.5%72.7%70.0%75.0%82.1%70.0%97.5%85.0%88.9%88.9%100.0%
Gemma 4 26B A4B QAT-Q4_0 💡 on* 🏠16.7%22.7%42.9%15.4%40.4%46.0%42.3%44.7%52.5%54.2%27.8%48.5%57.7%57.5%38.9%45.0%50.0%70.0%77.8%77.8%69.2%77.8%70.0%60.7%72.5%55.6%77.8%85.0%72.7%81.8%85.0%100.0%92.3%70.0%85.0%90.0%94.4%88.9%100.0%
Gemma 4 12B Q8_0 💡 on* 🏠15.0%27.3%40.0%33.3%36.0%34.8%39.1%34.2%52.0%27.3%50.0%47.7%42.3%52.5%66.7%54.2%57.1%67.5%72.2%83.3%77.3%88.9%52.5%75.0%60.5%66.7%83.3%72.5%69.2%86.4%61.1%94.7%83.3%94.7%85.0%87.5%100.0%100.0%100.0%
Laguna S 2.1 UD Q4_K_M 💡 on* 🏠19.0%15.4%32.5%54.2%30.0%37.5%25.0%40.0%47.5%59.1%42.5%45.0%42.5%47.5%65.0%52.5%58.3%35.0%60.0%80.0%72.7%90.0%57.5%72.7%65.0%77.8%88.9%55.0%100.0%86.4%77.8%87.5%90.9%75.0%90.0%75.0%95.0%100.0%90.0%
Gemma 4 26B A4B QAT-Q4_0 🏠0.0%31.8%9.1%22.2%18.2%31.8%25.0%54.5%25.0%59.1%25.0%44.4%61.1%33.3%35.0%66.7%18.2%70.0%65.0%80.0%40.9%75.0%72.2%45.5%75.0%85.0%60.0%80.0%86.4%72.7%66.7%88.9%63.6%100.0%90.0%100.0%90.0%90.0%100.0%
Gemma 4 12B Q8_0 🏠10.0%13.6%17.5%22.7%22.5%34.2%36.8%52.6%42.5%9.1%42.5%60.0%55.0%45.8%47.5%33.3%58.3%50.0%75.0%38.9%63.6%40.0%67.5%54.5%58.9%66.7%88.9%80.0%58.3%77.3%55.0%84.2%95.8%89.5%90.0%87.5%88.9%77.8%100.0%
Granite 4.2 30B Q8_0 💡 full 🏠25.0%16.7%14.3%25.0%19.4%32.4%39.3%30.8%42.3%26.9%30.8%35.7%50.0%42.9%41.7%81.8%41.7%45.8%50.0%45.5%62.5%60.0%58.3%77.3%50.0%90.0%63.6%50.0%73.1%64.3%45.5%83.3%90.9%57.1%81.8%64.3%73.1%70.0%70.0%
Nemotron 3 Super 120B A12B Q4_K_M 💡 full 🏠21.4%26.9%42.9%25.0%20.0%25.0%32.5%15.0%30.0%18.2%25.0%45.0%30.0%32.5%65.0%30.0%50.0%54.2%44.4%55.0%63.6%45.0%50.0%63.6%42.5%70.0%77.8%50.0%41.7%45.5%55.6%77.5%90.9%75.0%77.5%82.5%70.0%100.0%70.0%
Gemma 4 26B A4B Q8_0 🏠18.2%18.2%9.1%20.0%31.8%13.6%10.0%36.4%0.0%18.2%11.1%27.8%22.2%27.8%40.0%35.0%25.0%50.0%55.6%55.0%50.0%50.0%88.9%70.0%55.0%80.0%61.1%80.0%77.3%72.7%66.7%66.7%70.0%66.7%80.0%88.9%90.0%90.0%90.0%
Gemma 4 12B QAT-Q4_0 🏠0.0%27.3%9.1%9.1%9.1%22.7%20.0%13.6%22.2%27.3%44.4%33.3%22.2%16.7%20.0%20.0%61.1%54.5%45.0%45.0%50.0%50.0%66.7%60.0%80.0%45.0%60.0%77.8%54.5%68.2%65.0%70.0%65.0%50.0%100.0%77.8%77.8%70.0%70.0%
Granite 4.2 8B Q8_0 💡 full 🏠15.4%8.3%15.4%18.8%0.0%7.7%9.1%30.8%25.0%38.5%23.3%30.8%30.8%22.7%27.3%59.1%36.4%37.5%36.4%50.0%50.0%40.0%27.3%54.2%36.4%54.5%63.6%42.9%46.2%57.1%63.6%66.7%68.2%63.6%72.7%64.3%61.5%70.0%70.0%
Mistral Medium 3.5 💡 high0.0%8.3%8.3%0.0%16.7%13.6%27.4%22.7%44.4%25.0%22.2%27.8%22.2%11.1%10.0%25.0%60.0%40.0%55.0%50.0%50.0%60.0%30.0%50.0%60.0%50.0%55.0%77.8%54.2%79.2%55.6%50.0%50.0%55.6%70.0%100.0%70.0%80.0%100.0%
Mistral Medium 3.5 128B Q4_K_M 🏠17.5%0.0%16.7%29.2%26.0%22.5%22.5%20.0%27.5%27.3%33.3%42.5%30.0%47.5%42.5%27.8%32.5%41.7%50.0%11.1%33.3%72.7%70.0%66.7%55.0%45.0%66.7%50.0%53.8%50.0%33.3%70.0%58.3%65.0%60.0%57.5%70.0%55.6%44.4%
Granite 4.2 8B Q4_K_M 💡 full 🏠19.2%12.5%7.7%21.9%15.6%23.3%32.1%26.9%17.9%30.8%36.7%21.4%39.3%25.0%27.3%54.5%45.5%22.7%36.4%30.0%40.0%45.8%50.0%33.3%63.6%50.0%27.3%53.3%61.5%71.4%59.1%35.7%54.5%75.0%54.5%76.7%61.5%30.0%60.0%
Qwen3.6 35B A3B Q8_0 🏠5.0%0.0%10.0%36.4%7.5%23.7%12.5%10.5%32.5%13.6%12.5%30.0%27.5%39.5%35.0%25.0%41.1%50.0%57.5%45.0%20.0%63.6%40.0%45.0%36.4%40.0%60.4%70.0%45.5%54.5%75.0%76.3%72.7%84.2%70.0%70.0%80.0%80.0%90.0%
Qwen3.6 35B A3B Q4_K_M 🏠9.1%18.2%18.2%22.7%9.1%0.0%10.0%18.2%11.1%18.2%5.0%33.3%44.4%33.3%22.2%15.0%33.3%10.0%30.0%20.0%55.0%45.5%50.0%55.0%50.0%60.0%60.0%65.0%50.0%72.7%40.0%44.4%80.0%80.0%77.8%80.0%83.3%60.0%90.0%
Nemotron 3 Nano Omni 30B A3B Q8_0 💡 on* 🏠8.3%7.7%8.3%27.3%10.7%27.3%20.0%27.3%11.1%15.4%16.7%22.2%22.2%16.7%11.1%40.0%11.1%36.4%22.2%38.9%40.0%36.4%45.0%33.3%72.7%39.6%40.0%38.9%69.2%63.6%77.8%77.8%77.3%66.7%77.8%88.9%75.0%65.0%90.0%
Gemma 4 E2B Q8_0 💡 on* 🏠9.5%13.6%23.8%16.7%15.0%22.5%17.5%27.5%12.5%22.7%20.0%12.5%15.0%27.5%45.0%20.0%20.0%50.0%50.0%20.0%22.2%57.1%22.2%50.0%46.7%30.0%35.0%61.1%41.7%54.5%33.3%50.0%64.3%57.5%60.0%67.5%66.7%65.0%60.0%
Qwen3.8 27B Q8_0 🏠0.0%9.1%0.0%25.0%12.5%16.7%25.0%9.1%16.7%18.2%27.3%54.5%27.3%30.8%0.0%13.6%41.7%26.9%58.3%22.7%45.5%53.8%45.8%46.2%38.5%54.5%50.0%30.8%58.3%57.1%54.5%29.2%69.2%63.6%66.7%75.0%59.1%63.6%72.7%
Qwen3.8 27B Q4_K_M 🏠36.4%12.5%26.9%27.3%9.1%9.1%20.8%31.8%13.6%18.2%22.7%27.3%18.2%13.6%13.6%27.3%22.7%35.7%54.5%27.3%31.8%42.9%20.8%50.0%28.6%45.5%27.3%36.4%45.5%42.9%63.6%45.5%42.9%45.5%54.5%45.5%72.7%54.5%63.6%
Gemma 4 E4B Q4_K_M 🏠16.7%9.1%8.3%27.3%9.1%15.0%20.0%13.6%16.7%9.1%16.7%30.0%15.0%38.9%22.2%33.3%45.0%54.5%44.4%33.3%35.0%36.4%44.4%66.7%40.9%25.0%60.0%22.2%66.7%45.5%36.4%60.0%40.9%44.4%77.8%77.8%70.0%44.4%55.6%
Gemma 4 E4B Q8_0 🏠15.0%9.1%22.5%31.8%0.0%7.9%10.5%10.5%10.0%22.7%5.0%25.0%0.0%5.3%12.5%11.1%15.8%16.7%22.5%33.3%30.0%33.3%50.0%30.0%64.3%23.7%55.6%22.2%50.0%70.8%54.5%40.0%57.1%37.5%45.0%62.5%85.0%50.0%85.0%
Granite 4.2 30B Q4_K_M 💡 full 🏠0.0%0.0%15.4%9.4%6.7%13.3%3.6%7.7%14.3%16.7%7.1%17.9%7.7%16.7%9.1%36.4%4.2%9.1%9.1%30.0%35.0%31.8%50.0%41.7%45.5%27.3%20.0%22.7%35.7%30.8%57.1%59.1%42.9%40.9%45.5%60.7%46.4%65.0%50.0%
Gemma 4 E4B Q8_0 💡 on* 🏠0.0%9.1%5.0%0.0%0.0%21.1%10.5%0.0%5.0%18.2%0.0%30.0%30.0%5.3%25.0%0.0%10.5%42.9%25.0%33.3%50.0%36.4%44.4%35.0%25.0%15.8%20.0%33.3%42.5%36.4%54.5%55.6%62.5%59.1%50.0%65.0%44.4%40.0%66.7%
Nemotron 3 Nano Omni 30B A3B Q4_K_M 💡 on* 🏠0.0%11.5%5.0%8.3%10.0%10.0%5.0%5.0%20.0%9.1%10.0%2.5%15.0%15.0%10.0%10.0%10.0%18.2%22.5%20.0%0.0%27.3%30.0%40.0%45.5%30.0%22.2%22.2%40.0%33.3%45.5%22.2%55.0%54.5%50.0%60.0%60.0%44.4%40.0%
Gemma 4 E2B Q4_K_M 💡 on* 🏠4.8%0.0%19.0%0.0%5.0%12.5%22.5%15.0%0.0%0.0%5.0%15.0%10.0%12.5%25.0%0.0%12.5%35.7%17.5%11.1%22.2%35.7%0.0%42.5%23.3%30.0%20.0%11.1%32.5%25.0%54.5%22.2%37.5%39.3%35.0%40.0%44.4%55.6%35.0%
Gemma 4 E4B Q4_K_M 💡 on* 🏠16.7%0.0%25.0%0.0%0.0%0.0%0.0%9.1%0.0%4.5%0.0%11.1%5.6%0.0%5.0%10.0%11.1%26.9%30.0%10.0%22.2%38.5%30.0%30.0%38.5%20.0%16.7%25.0%33.3%40.9%27.3%30.0%15.0%53.6%55.6%40.0%55.6%55.0%60.0%
Gemma 4 E2B Q8_0 🏠9.1%9.1%18.2%9.1%0.0%0.0%0.0%0.0%5.6%0.0%16.7%11.1%11.1%0.0%0.0%10.0%22.2%30.0%0.0%10.0%30.0%30.0%20.0%44.4%70.0%20.0%40.0%35.0%35.0%36.4%45.5%55.6%50.0%35.0%60.0%55.6%44.4%45.0%50.0%
Gemma 4 E2B Q4_K_M 🏠9.1%0.0%9.1%0.0%0.0%9.1%10.0%0.0%0.0%0.0%10.0%0.0%0.0%0.0%10.0%0.0%0.0%30.0%30.0%10.0%30.0%30.0%0.0%55.6%40.0%10.0%10.0%10.0%40.0%27.3%36.4%44.4%15.0%50.0%33.3%60.0%65.0%40.0%50.0%

Published duel coverage

Each cell shows how many published AI-judged duels across all published evaluation protocols compare the two models for this language and category. The models on both axes are candidates; this matrix does not show which judge produced each verdict.

Symmetric matrix of published duel counts. The diagonal is zero because a model is not compared with itself.

💡 marks a model with reasoning, ↯ one without reasoning. 🏠 marks a model run locally on the site author’s computer. “on*” means reasoning is enabled or disabled (local models with no intermediate levels): the requested fine-grained setting is not supported by the local file and falls back to the default enabled mode. For local models with adjustable reasoning (Qwen3.8, Muse Glimmer, Granite 4.2, Nemotron 3 Super), the displayed level (xhigh**, high, full) is the model’s default level: the requested fine-grained setting is not supported by the local file, which falls back to its default behavior.

View as a data table
Published duel counts by model pair
ModelMuse Glimmer 30B K-Quant Dynamic 💡 high 🏠Qwen3.8 27B Q4_K_M 💡 xhigh** 🏠Muse Glimmer 30B K-Quant 17GB 💡 high 🏠Qwen3.8 27B Q8_0 💡 xhigh** 🏠Gemma 4 31B Q8_0 💡 on* 🏠Gemma 4 31B QAT-Q4_0 💡 on* 🏠Qwen3.6 35B A3B Q8_0 💡 on* 🏠Gemma 4 31B QAT-Q4_0 🏠Gemma 4 12B QAT-Q4_0 💡 on* 🏠Gemma 4 31B Q8_0 🏠Gemma 4 26B A4B Q8_0 💡 on* 🏠Qwen3.6 35B A3B Q4_K_M 💡 on* 🏠Gemma 4 26B A4B QAT-Q4_0 💡 on* 🏠Gemma 4 12B Q8_0 💡 on* 🏠Laguna S 2.1 UD Q4_K_M 💡 on* 🏠Gemma 4 26B A4B QAT-Q4_0 🏠Gemma 4 12B Q8_0 🏠Granite 4.2 30B Q8_0 💡 full 🏠Nemotron 3 Super 120B A12B Q4_K_M 💡 full 🏠Gemma 4 26B A4B Q8_0 🏠Gemma 4 12B QAT-Q4_0 🏠Granite 4.2 8B Q8_0 💡 full 🏠Mistral Medium 3.5 💡 highMistral Medium 3.5 128B Q4_K_M 🏠Granite 4.2 8B Q4_K_M 💡 full 🏠Qwen3.6 35B A3B Q8_0 🏠Qwen3.6 35B A3B Q4_K_M 🏠Nemotron 3 Nano Omni 30B A3B Q8_0 💡 on* 🏠Gemma 4 E2B Q8_0 💡 on* 🏠Qwen3.8 27B Q8_0 🏠Qwen3.8 27B Q4_K_M 🏠Gemma 4 E4B Q4_K_M 🏠Gemma 4 E4B Q8_0 🏠Granite 4.2 30B Q4_K_M 💡 full 🏠Gemma 4 E4B Q8_0 💡 on* 🏠Nemotron 3 Nano Omni 30B A3B Q4_K_M 💡 on* 🏠Gemma 4 E2B Q4_K_M 💡 on* 🏠Gemma 4 E4B Q4_K_M 💡 on* 🏠Gemma 4 E2B Q8_0 🏠Gemma 4 E2B Q4_K_M 🏠
Muse Glimmer 30B K-Quant Dynamic 💡 high 🏠0132612202020202011202121202111201421111113122013201112211211122013202021121111
Qwen3.8 27B Q4_K_M 💡 xhigh** 🏠1301311111111111111111111111311111213111112121112111113111112111112111311111111
Muse Glimmer 30B K-Quant 17GB 💡 high 🏠2613012212020202011202121202011201421111113122113201112211213122013202021121111
Qwen3.8 27B Q8_0 💡 xhigh** 🏠1211120111111111211111113121227111612201116121216111111121211111116121212111111
Gemma 4 31B Q8_0 💡 on* 🏠2011211102825202912352826252011201820111115122516201114201211112015202020111111
Gemma 4 31B QAT-Q4_0 💡 on* 🏠2011201128024202611246262232011191720111113112015191111201211101915192020101111
Qwen3.6 35B A3B Q8_0 💡 on* 🏠2011201125240192511202326232010191420101011312014241010201212151914192020101010
Gemma 4 31B QAT-Q4_0 🏠2011201120201902011201919192011191320111113112013191111201111111913192020111111
Gemma 4 12B QAT-Q4_0 💡 on* 🏠201120122926252001223292025201020132010914920142099201211920142020209910
Gemma 4 31B Q8_0 🏠1111111112111111120111212111111111311111113121113111113111111111112111111111111
Gemma 4 26B A4B Q8_0 💡 on* 🏠201120113524202023110262720201020132099159241520109201111920142020209910
Qwen3.6 35B A3B Q4_K_M 💡 on* 🏠21112111286223192912260682220920142099139201420992011111020142020209910
Gemma 4 26B A4B QAT-Q4_0 💡 on* 🏠21112113266226192012276802620920142099139201420992011111020132020209910
Gemma 4 12B Q8_0 💡 on* 🏠201120122523231925112022260209241420991192014199920131191915192020999
Laguna S 2.1 UD Q4_K_M 💡 on* 🏠211320122020202020112020202001020122010101110201120992012119201120202010910
Gemma 4 26B A4B QAT-Q4_0 🏠11111127111110111011109991009111010101110911101010101111991110109101010
Gemma 4 12B Q8_0 🏠201120112019191920112020202420901220109111020112899201211101912192020999
Granite 4.2 30B Q8_0 💡 full 🏠1412141618171413131313141414121112012111112101211121011141314111211141114131010
Nemotron 3 Super 120B A12B Q4_K_M 💡 full 🏠2113211220202020201120202020201020120910111020112010920121192011202020101010
Gemma 4 26B A4B Q8_0 🏠11111120111110111011999910101011901010109101010910111199109109101010
Gemma 4 12B QAT-Q4_0 🏠111111111111101191199991010911101001010910101010911111010101010991010
Granite 4.2 8B Q8_0 💡 full 🏠1312131615131113141315131311111111121110100101112111111141314111211111114131010
Mistral Medium 3.5 💡 high1212121212113111912999910101010101010100101010101091212910109109101010
Mistral Medium 3.5 128B Q4_K_M 🏠201121122520202020112420202020920122099111001220109201313920122020201099
Granite 4.2 8B Q4_K_M 💡 full 🏠1312131616151413141315141414111111111110101210120111011151314111411121115131010
Qwen3.6 35B A3B Q8_0 🏠2011201120192419201120202019201028122010101110201101048201111101911192020101010
Qwen3.6 35B A3B Q4_K_M 🏠1111111111111011911109999109101010101110101010010101611109101091091010
Nemotron 3 Nano Omni 30B A3B Q8_0 💡 on* 🏠121312111411101191399999109119910111091148100913119911999101010
Gemma 4 E2B Q8_0 💡 on* 🏠21112112202020202011202020202010201420109149201520109012119201420202091010
Qwen3.8 27B Q8_0 🏠1211121212121211121111111113121112131211111312131311161312014111213111212111111
Qwen3.8 27B Q4_K_M 🏠1112131111111211111111111111111111141111111412131411111111140111114111111111111
Gemma 4 E4B Q4_K_M 🏠121112111110151191191010999101199101199111010991111010119991099
Gemma 4 E4B Q8_0 🏠201120112019191920112020201920919122091012102014199920121110014202020101010
Granite 4.2 30B Q4_K_M 💡 full 🏠1312131615151413141214141315111112111110101110121111101114131411140111114141010
Gemma 4 E4B Q8_0 💡 on* 🏠2011201220191919201120202019201019142091011920121910920111192011020209109
Nemotron 3 Nano Omni 30B A3B Q4_K_M 💡 on* 🏠201320122020202020112020202020102011201010111020112099201211920112002010910
Gemma 4 E2B Q4_K_M 💡 on* 🏠21112112202020202011202020202092014209914920152010920121192014202009910
Gemma 4 E4B Q4_K_M 💡 on* 🏠12111211111010119119999101091310109131010131091091111101014910901010
Gemma 4 E2B Q8_0 🏠11111111111110119119999910910101010101091010101010111191010109910010
Gemma 4 E2B Q4_K_M 🏠11111111111110111011101010910109101010101010910101010101111910109101010100

Which local AI writes best?

If you want one recommendation rather than a methodology, this is it:

What counts as "local" here

Four conditions, applied strictly:

  1. Open weights. The model files must be publicly downloadable and runnable outside the vendor's infrastructure.
  2. Genuinely runnable offline. Not a gated API with an open-sounding name.
  3. Under 256 billion total parameters. Strictly below the threshold.
  4. Already evaluated on this site. No estimates, no placeholder entries in the ranking itself.

The parameter limit applies to total parameters, including for mixture-of-experts models. This is worth stating plainly because it is a common and consequential error: an MoE model that activates a small fraction of its parameters per token still has to have all of its weights in memory. Total parameters determine eligibility and memory requirements; active parameters affect speed. When both counts are known, the model name keeps both (for example, 35B A3B) so the active count can inform speed expectations without being mistaken for a memory figure.

One thing the criteria do not require: that we generated the model's responses on a machine in this office. What makes a model eligible is that the weights are available and the model can be run locally — not the historical location of the hardware that produced its texts. Where a variant's throughput was measured on our own hardware, the speed figure is labelled as an AIWB measurement. Figures taken from elsewhere are labelled as external or estimated. We would rather show a labelled estimate than a confident invention.

Open source is not the same as open weights

"Open source" is the phrase people search for, so it is worth untangling. Open source is a licensing claim broader than merely publishing weights; how that framework applies to all model artefacts is still contested. Most models in this ranking publish their weights under licences that range from permissive to distinctly restrictive — commercial-use limits, acceptable-use clauses, redistribution conditions.

We rank on writing quality. The licence is a separate question, and an important one if you are shipping a product rather than drafting an email. Check it before you build on a model; it has no bearing on how well the model writes.

Eligible, but not yet tested

Models appear here only once they have been evaluated. Anything that meets the criteria but hasn't been through the corpus yet sits in a short waiting list, without a score and without a rank, so you can see what's coming rather than wonder whether we've noticed it.

Quantization, without the marketing

Quantization stores a model's weights using fewer bits per value. A model trained in 16-bit precision can be repackaged at 8 bits, or 4, shrinking the download and, more importantly, the memory footprint. On the kind of inference most people run at home — one user, short batches, bottlenecked by how fast memory can be read rather than by raw compute — a smaller file often runs faster too, because there is simply less data to move per token.

It also changes the model. Sometimes barely; sometimes visibly.

What we can say with confidence is that the size of the effect is not a constant. It varies with the model, the quantization method, the task, the language, the context length and the level of quantization. This is why you will not find a "Q4 keeps 98% of quality" claim on this page. That number does not exist as a general fact, and repeating it is how people end up disappointed by a specific model on a specific task.

The tags you'll see in the filter:

  • Q8 — an eight-bit variant. We treat it as the practical high-quality reference point in this ranking: it is the version we reach for when memory allows. That is a working convention, not a claim that Q8 is mathematically equivalent to BF16 or FP16. Whether the remaining gap is measurable in English prose is an empirical question, and we prefer to answer it with paired tests rather than assertion.
  • Q4 QAT — quantization-aware training. The model was prepared or trained with four-bit execution in mind, which makes it a distinct checkpoint rather than a post-hoc compression of an existing one.
  • Q4_K_M — a four-bit GGUF quantization applied after training, with mixed precision across different parts of the model. This is not QAT, and conflating the two is a genuine source of confusion. Included by default.
  • Q6 — a middle ground that appears when an evaluated variant exists.

Filters apply as you change them, and the state is carried in the URL, so a filtered view can be linked or bookmarked.

Muse Glimmer's two K-Quant builds

Meta publishes Muse Glimmer 30B in two unusual GGUF builds that average approximately four bits per weight. K-Quant Dynamic keeps selected tensors at higher precision and targets 32 GB of VRAM; K-Quant 17GB compresses more aggressively and targets 24 GB. They are not Q4_K_M files, so we keep their exact names and score them as separate variants. Both sit under the broad Q4 filter because that filter groups practical four-bit-class downloads rather than claiming an identical encoding.

Meta reports average degradation of 0.2% and 1.0% respectively across fifteen common benchmarks. Those are publisher measurements, not AIWB writing scores and not guarantees for English, French, German or Spanish prose. Our rankings and direct comparison measure the two files independently on native writing prompts.

Hardware recommendations

8 GB

Best writing quality

Granite 4.2 8B Q4_K_M 💡 full 🏠

24.6 · local rank 36

Faster measured alternative

Gemma 4 E2B Q4_K_M 💡 on* 🏠

39.7 tok/s · AIWB measurement on Strix Halo

16 GB

Best writing quality

Gemma 4 12B QAT-Q4_0 💡 on* 🏠

46.8 · local rank 16

Faster measured alternative

Gemma 4 E2B Q4_K_M 💡 on* 🏠

39.7 tok/s · AIWB measurement on Strix Halo

32 GB

Best writing quality

Muse Glimmer 30B K-Quant Dynamic 💡 high 🏠

76.8 · local rank 1

Faster measured alternative

Qwen3.6 35B A3B Q4_K_M 💡 on* 🏠

71.2 tok/s · AIWB measurement on Strix Halo

Loading guidance, not guarantees: context, KV cache, batch and runtime change the footprint. AIWB speeds were measured on the Strix Halo machine described on the About page and do not predict another platform.

Why bother running it locally at all

The cloud models at the top of our general leaderboard are strong. Local models are chosen for reasons that have nothing to do with beating them.

When the runtime and its integrations stay local, your text need not leave the machine — which is the whole argument for anyone drafting confidential, legal, medical or unpublished material. It works with no connection, on a plane or a bad hotel connection. It doesn't change underneath you: a checkpoint you downloaded behaves the same next month, where a cloud endpoint can be updated, deprecated or retired without warning. There's no per-request cost and no rate limit. And you can inspect, fine-tune and integrate it on your own terms.

Set against that: you are trading some quality for that control, and you're providing the hardware. Whether the trade is worth it is exactly the question this ranking exists to inform.

Quality against speed

These are separate axes, and we keep them separate on purpose. The ranking column measures how well a variant writes English. The speed figures measure how fast it produces tokens on specific hardware. Nothing in this page combines them into a single "best" score, because the right balance depends on what you're doing.

Drafting a long article, where you'll read and edit the output carefully, rewards quality. Interactive back-and-forth — rephrasing, brainstorming, dialogue — rewards responsiveness, and a slightly weaker model that answers in two seconds may serve you better than a stronger one that takes thirty. Both variants are listed in every hardware tier for that reason.

FAQ

What is the best local AI for writing in English? The live summary above names the current top-ranked variant with its score and sample count. Bear in mind that the top variant is only the best answer if it fits your memory with room for your context; the per-tier recommendations above are more useful for most readers.

What does "local LLM" mean? An LLM is a large language model — the type of AI that generates text. A local LLM is one you download and run on your own computer, offline, rather than calling a service over the internet.

Why is there a 256B parameter limit? The threshold keeps this page centred on models that remain plausible on personal or compact workstation hardware. It applies to total parameters, including for mixture-of-experts models, because all the weights must be in memory even when only a fraction is active per token. Larger systems are outside this page's scope rather than impossible in absolute terms.

Does quantization ruin quality? Not inherently, and not by a fixed amount. The effect depends on the model, the method, the task and the language. We show quantizations as separate ranked rows so you can see the difference where we've measured it, instead of trusting a general percentage.

What's the difference between Q4 QAT and Q4_K_M? Q4 QAT is a distinct checkpoint prepared or trained for four-bit execution. Q4_K_M is a four-bit GGUF quantization applied to an existing model after training, using mixed precision across its layers. They are not interchangeable and we never merge them into one row.

How much VRAM do I need? More than the model file. Budget for the weights plus the runtime plus the KV cache, which grows with your context length. The dedicated GPU tab gives recommendations for common VRAM sizes with realistic headroom built in.

Is unified memory as good as VRAM? For capacity, often better — unified-memory machines can hold models no consumer graphics card can. For bandwidth, generally not, and that difference shows up as slower generation. Don't treat the two as equivalent.

Are open-weight models open source? Not automatically. Publishing weights does not by itself establish an open-source licence, and model licences vary considerably in what they permit commercially. The licence is independent of writing quality — check it separately.

Why do these scores include comparisons against cloud models? Because it keeps every variant on one scale and preserves its full match history. The rank is local; the score comes from the general English leaderboard.

How often does this page update? Continuously. Rankings, filters and hardware recommendations are generated from live data, and new variants enter the ranking as soon as they've been evaluated on the English corpus.