Methodology

How AIWB collects and reports writing preferences.

Native, language-specific tasks

Each comparison uses one reviewed writing prompt and two pre-generated responses in the selected language. No model runs on the public server.

Blind human comparisons

Model names are hidden while a visitor compares response A with response B. The display order is stable for that visitor and recorded with the vote.

Votes and exclusions

A better, B better, tie, and both bad count toward the human results. Bug or invalid and skip are recorded but excluded; both bad gives both models a loss.

Provisional, separated results

AI-judged pairwise duels determine the provisional main ranking. Human votes form an independent signal; the two channels are reported separately and never merged.