Fiction Eval leaderboard · model review
Claude Opus 4.8 for fiction writing — blind-judged review
July 2026: 10 fresh genre scenes, three tries each, judged blind against 7 other frontier models by a different AI family.
Where it lands
#1 of 8 — score 1927 (likely 1888–1976), record 150–14–43 across 207 blind matchups. Slop 14.88 per 1,000 words (3rd-cleanest of the field — a separate AI-tell count, not the quality ranking).
| # | Model | Score | Record | Slop |
|---|---|---|---|---|
| 1 | Claude Opus 4.8 | 1927 (likely 1888–1976) | 150–14–43 | 14.88 |
| 2 | Claude Sonnet 4.6 | 1914 (likely 1874–1961) | 146–15–46 | 18.57 |
| 3 | Grok 4.3 | 1856 (likely 1815–1899) | 137–28–42 | 9.43 |
| 4 | GLM 5.2 | 1776 (likely 1732–1822) | 117–42–42 | 17.53 |
| 5 | Gemini 3.1 Pro | 1471 (likely 1423–1519) | 78–104–13 | 13.45 |
| 6 | DeepSeek V3.2 | 1251 (likely 1201–1302) | 52–144–11 | 23.24 |
| 7 | Mistral Large | 1118 (likely 1054–1167) | 35–162–10 | 17.79 |
| 8 | Llama 4 Maverick | 686 (likely 627–741) | 0–206–1 | 22.92 |
Judged strengths and weaknesses
Traits where Claude Opus 4.8 scores above the field average: voice, pacing, dialogue, genre convention, avoiding AI tells. Below the field average: none. (Each rated 1–5 by the judge; field average taken across all eight models.)
Sample
LitRPG — “dungeon-core scene” scene; full version on the board.
Head-to-heads
vs Claude Sonnet 4.6 · vs Grok 4.3 · vs GLM 5.2 · vs Gemini 3.1 Pro · vs DeepSeek V3.2 · vs Mistral Large · vs Llama 4 Maverick
Writing a book of your own? Book Factory runs the same craft checks on full manuscripts.