What changed on AI Writing Benchmark

A dated record of public improvements, the checks behind them and the two-model consensus required before release.

A safer eight-hour audit cycle

The benchmark can now improve on a regular schedule without turning maintenance into an unchecked automation loop.

  • Medium: Added a bounded systemd audit every eight hours. GPT-5.6 Sol may implement only changes approved independently by Claude Opus 5.
  • Small: Added release gates that require restoration of the previous version after failed tests, public checks or a rejected post-publication review.
  • Small: Published this timestamped changelog in four independently written editions.

Checks: Full Django test suite, four-site public checks and independent Opus 5 review required.