Fiction Eval leaderboard · model review

GLM 5.2 for fiction writing — blind-judged review

July 2026: 10 fresh genre scenes, three tries each, judged blind against 7 other frontier models by a different AI family.

Where it lands

#4 of 8 — score 1776 (likely 1732–1822), record 117–42–42 across 201 blind matchups. Slop 17.53 per 1,000 words (4th-cleanest of the field — a separate AI-tell count, not the quality ranking).

#ModelScoreRecordSlop
1Claude Opus 4.81927 (likely 1888–1976)150–14–4314.88
2Claude Sonnet 4.61914 (likely 1874–1961)146–15–4618.57
3Grok 4.31856 (likely 1815–1899)137–28–429.43
4GLM 5.21776 (likely 1732–1822)117–42–4217.53
5Gemini 3.1 Pro1471 (likely 1423–1519)78–104–1313.45
6DeepSeek V3.21251 (likely 1201–1302)52–144–1123.24
7Mistral Large1118 (likely 1054–1167)35–162–1017.79
8Llama 4 Maverick686 (likely 627–741)0–206–122.92

Judged strengths and weaknesses

Traits where GLM 5.2 scores above the field average: voice, pacing, dialogue, genre convention, avoiding AI tells. Below the field average: none. (Each rated 1–5 by the judge; field average taken across all eight models.)

Sample

Romance — “romance turn” scene; full version on the board.

# The Inventory "You're counting wrong again." "I'm counting *correctly*. You're reading the sheet wrong." Mara shoved the clipboard at him. Dev took it, scanned the numbers, and shoved it back. They'd been in the stockroom for two hours. Two hours of fluorescent hum and dust and Dev's shoulder bumping hers every time he reached for the next shelf, which he was doing on purpose, she was sure of it, because he did everything on purpose. "Row seventeen," she said. "Fourteen units." "Twelve." "Dev. I can *see* fourteen boxes." "I can see fifteen. You miscounted." She turned to face him. He was already facing her, because of course he was. His eyes dropped to the clipboard in her hands, then back up. "Write down fourteen," she said. "Write down fifteen." "You're impossible." "You're stubborn." Those words had weight. They always had weight, between them. In the courtroom they landed like blows — her client's interests against his, his arguments against hers, each one sharpened to cut. Here, in the stockroom, …

Head-to-heads

vs Claude Opus 4.8 · vs Claude Sonnet 4.6 · vs Grok 4.3 · vs Gemini 3.1 Pro · vs DeepSeek V3.2 · vs Mistral Large · vs Llama 4 Maverick

How it works: every pair of models is judged blind on the same scene, with the passages' order flipped so being shown first can't sway it, by GPT-5.4 — a family that isn't on the board, so nobody scores their own side. Each score carries a likely range; overlapping ranges are called a tie. Slop is scored separately by a fixed checklist, not an AI. Full board, every prompt, and the FAQ: thebookfactoryai.com/board. Get each new board by email on the model-drop list.

Writing a book of your own? Book Factory runs the same craft checks on full manuscripts.