Fiction Eval leaderboard · model review
Gemini 3.1 Pro for fiction writing — blind-judged review
July 2026: 10 fresh genre scenes, three tries each, judged blind against 7 other frontier models by a different AI family.
Where it lands
#5 of 8 — score 1471 (likely 1423–1519), record 78–104–13 across 195 blind matchups. Slop 13.45 per 1,000 words (2nd-cleanest of the field — a separate AI-tell count, not the quality ranking).
| # | Model | Score | Record | Slop |
|---|---|---|---|---|
| 1 | Claude Opus 4.8 | 1927 (likely 1888–1976) | 150–14–43 | 14.88 |
| 2 | Claude Sonnet 4.6 | 1914 (likely 1874–1961) | 146–15–46 | 18.57 |
| 3 | Grok 4.3 | 1856 (likely 1815–1899) | 137–28–42 | 9.43 |
| 4 | GLM 5.2 | 1776 (likely 1732–1822) | 117–42–42 | 17.53 |
| 5 | Gemini 3.1 Pro | 1471 (likely 1423–1519) | 78–104–13 | 13.45 |
| 6 | DeepSeek V3.2 | 1251 (likely 1201–1302) | 52–144–11 | 23.24 |
| 7 | Mistral Large | 1118 (likely 1054–1167) | 35–162–10 | 17.79 |
| 8 | Llama 4 Maverick | 686 (likely 627–741) | 0–206–1 | 22.92 |
Judged strengths and weaknesses
Traits where Gemini 3.1 Pro scores above the field average: dialogue, genre convention. Below the field average: voice, pacing, avoiding AI tells. (Each rated 1–5 by the judge; field average taken across all eight models.)
Sample
Thriller — “thriller cold open” scene; full version on the board.
Head-to-heads
vs Claude Opus 4.8 · vs Claude Sonnet 4.6 · vs Grok 4.3 · vs GLM 5.2 · vs DeepSeek V3.2 · vs Mistral Large · vs Llama 4 Maverick
Writing a book of your own? Book Factory runs the same craft checks on full manuscripts.