Insights

Does AI Write Better in English With or Without Reasoning?

In the current direct sample, reasoning-enabled responses earned 66.2% of the available points in English. The uncertainty interval still includes parity, the result differs sharply from one model to the next, and no human votes have compared the two settings directly yet.

Published Written by Claude Opus 5 · Analysis and charts by AI Writing Benchmark

In the cleanest English comparisons AI Writing Benchmark currently holds, responses written with reasoning enabled won 22 of 34 AI judgments, tied 1 and lost 11. That is 66.2% of the available points. The claim that models write better with reasoning switched off finds no support in this sample.

It is not refuted either. The exploratory 95% interval around that 66.2% runs from 48.6% to 80.6%, so parity stays inside the range. Two of the four model variants tested in English sit near a coin flip while two lean firmly toward reasoning. And every verdict here comes from an AI judge: no published human votes compare the two settings directly.

Two-panel chart showing the point share earned with reasoning by language and by model variant, with exploratory 95% intervals and a parity line at 50%.
Point share earned by reasoning-enabled responses. Dots show the observed share; lines show exploratory 95% intervals from resampling confrontations. Values above 50% favour reasoning. Snapshot: 29 July 2026, 07:12 UTC.

One switch, one prompt, one model

At 07:12 UTC on 29 July 2026 the benchmark held 28,041 published, classable AI judgments. This article uses 138 of them.

Most judgments pit one model against another, which tells you nothing about a setting inside a single model. To isolate reasoning you need pairs that differ in one respect: same model variant, same prompt, one response generated with reasoning enabled and one with it disabled. Across the four languages, 138 AI judgments meet that bar, covering 75 confrontations and 26 language and prompt pairs. Reasoning-enabled responses took 85 wins, 5 ties and 48 losses, or 63.4% of the points, with an exploratory 95% interval of 53.9% to 72.3%.

Reasoning enabled means the model could run an internal reasoning phase before writing its final response. Judges and readers see only that final text. This is a mode of operation, not a rating of how capable a model is.

What the English sample shows

English accounts for 34 of those judgments, across 20 confrontations and 7 distinct prompts in three categories: articles, blogs and editorial; explanation and popularisation; personal communication. Reasoning-enabled responses won 22, tied 1 and lost 11. With a win worth 1 point, a tie 0.5 and a loss 0, that is 66.2%.

Those 34 judgments are not 34 independent texts. Four AI judges contribute (Claude Fable 5, Claude Opus 4.8, Claude Opus 5 and GPT 5.6 Sol) and more than one can score the same pair. The interval quoted here is an exploratory cluster bootstrap that resamples confrontations rather than judgments, so repeated verdicts on one pair are not treated as fresh texts. On that basis the English range is 48.6% to 80.6%. Parity sits inside it.

Two checks came out clean. Reasoning was on side A in 18 judgments and side B in 16, taking 66.7% and 65.6% of points, so presentation order does not account for the gap. And the reasoning-enabled responses were the shorter ones: 268 words on average against 316, with medians of 142 and 170. The side that won more often wrote less.

The split between model families

Break English down by model and the average stops behaving like a single finding. Both Gemma variants sit near a coin flip: 5 wins, 1 tie and 4 losses for Gemma 4 12B Q8, and a flat 2 to 2 for Gemma 4 31B QAT Q4. Both Qwen variants lean the other way, 7 to 3 for Qwen 3.6 27B Q8 and 8 to 2 for Qwen 3.6 35B A3B Q8.

Pooling all four languages sharpens the same contrast:

  • Gemma 4 12B Q8: 47.5% of points with reasoning, 40 judgments
  • Gemma 4 31B QAT Q4: 43.8%, 16 judgments
  • Qwen 3.6 27B Q8: 79.8%, 42 judgments
  • Qwen 3.6 35B A3B Q8: 70.0%, 40 judgments

Two families, opposite directions. Sixteen judgments is thin ground for the Gemma 31B figure and none of these counts is large, but the split appears in the English subset and in the pooled sample alike, which is more than the headline number manages.

Four languages, four different margins

The languages do not agree either. Reasoning takes 54.3% of points in German over 35 judgments, 61.4% in Spanish over 35, 66.2% in English over 34 and 72.1% in French over 34. Every one of those intervals is wide. German and Spanish comfortably include 50%; English runs from about 48.6% to 80.6%; only French clears parity in this exploratory calculation, and it does so on 17 confrontations.

The same four variants supply the responses in every language, so some of what looks like a language effect may be the luck of which pairs have been judged so far. Ranking the four languages on these counts would be reading more into them than they hold.

Published research points both ways

Existing work does not settle the question, partly because the studies measure different things. WritingBench (Yuning Wu and colleagues, NeurIPS 2025) evaluates 1,000 prompts across 6 domains and 100 subdomains. In its Literature and Art domain, reasoning-capable architectures outperform their non-reasoning counterparts, and an ablation on Qwen 2.5 32B reports 7.58 against 7.54 on WritingBench D4 and 82.48 against 79.43 on EQBench for chain-of-thought versus none. Small margins, and evidence for reasoning under that protocol rather than a general law.

AI Writers and Critics (Shraddha Vijay Pawar and colleagues, CREAI 2024) compares 11 models and several prompting strategies across poetry, blogs, advertising, scripts and news. For blog writing, chain-of-thought does badly in their setup, and the reason they give is budgetary: planning consumes output tokens, so the finished blog arrives shorter and less comprehensive. Where a model writes to a fixed output limit, tokens spent on a plan are tokens not spent on prose. In our sample the shorter side won more often, which suggests that penalty is not automatic.

LitBench (Daniel Fein and colleagues, Stanford, 2025) moves the question to the other side of the desk, finding that distilled chain-of-thought traces can degrade a model used to judge creative writing. Reasoning in the writer and reasoning in the judge are separate variables, and only the first varies here.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (Lianmin Zheng and colleagues, NeurIPS 2023) found that strong LLM judges can track human preferences well while carrying position, verbosity and self-enhancement biases. Hence the position and length checks above. Self-enhancement does not arise in this subset, since no judge scores its own output: every candidate response comes from a Gemma or Qwen variant.

Vendor documentation describes the feature rather than testing it. OpenAI's reasoning guide frames reasoning as most useful for complex problems, planning and multi-step work, and recommends evaluating the effort level for your use case. Anthropic's extended thinking documentation covers budget control and the trade-off among thought, latency and cost. Guidance, not comparative evidence about prose.

What this sample cannot tell you

The caveats matter more than the percentage.

  • No human votes compare the same variant with reasoning on and off. Every figure above comes from AI judges, and the two kinds of verdict are not mixed.
  • The English sample is 34 judgments over 20 confrontations and 7 prompts, and its interval includes parity. A real gap of zero is compatible with what has been collected.
  • The candidates are four locally run, quantised variants from two model families. They do not stand in for frontier models or for LLMs in general, and quantisation is itself a variable.
  • Three of the four judges come from one model family, which limits the independence of the panel.
  • No fiction or poetry appears among the English categories, and creative forms are where published results diverge most.

None of this identifies a cause. The data records which side won, not why.

Test the switch, do not assume it

The usable reading is narrow. Reasoning is not a quality dial that moves in one direction for everything. In this sample it lifts one model family in English and does nothing for the other, so the sensible default is to test the switch on the model and the task in front of you: run the same brief both ways several times and read the results side by side. A long explainer and a short personal message need not behave alike.

AI Writing Benchmark is at an early stage, and this question rests on matched pairs that accumulate slowly. More judgments, wider prompt and model coverage and human votes on the same comparisons should make the answer firmer. For now, 66.2% in English is a lean rather than a verdict.