Fiction Eval leaderboard · model review

Claude Sonnet 4.6 for fiction writing — blind-judged review

July 2026: 10 fresh genre scenes, three tries each, judged blind against 7 other frontier models by a different AI family.

Where it lands

#2 of 8 — score 1914 (likely 1874–1961), record 146–15–46 across 207 blind matchups. Slop 18.57 per 1,000 words (6th-cleanest of the field — a separate AI-tell count, not the quality ranking).

#ModelScoreRecordSlop
1Claude Opus 4.81927 (likely 1888–1976)150–14–4314.88
2Claude Sonnet 4.61914 (likely 1874–1961)146–15–4618.57
3Grok 4.31856 (likely 1815–1899)137–28–429.43
4GLM 5.21776 (likely 1732–1822)117–42–4217.53
5Gemini 3.1 Pro1471 (likely 1423–1519)78–104–1313.45
6DeepSeek V3.21251 (likely 1201–1302)52–144–1123.24
7Mistral Large1118 (likely 1054–1167)35–162–1017.79
8Llama 4 Maverick686 (likely 627–741)0–206–122.92

Judged strengths and weaknesses

Traits where Claude Sonnet 4.6 scores above the field average: voice, pacing, dialogue, genre convention, avoiding AI tells. Below the field average: none. (Each rated 1–5 by the judge; field average taken across all eight models.)

Sample

Progression fantasy — “progression beat” scene; full version on the board.

The elder's laughter was the worst of it. Not the public correction, not the demonstration performed on Kael's body like he was practice equipment, not even the forty students watching from the courtyard's edges. The laughter—bemused, almost fond, the way you'd laugh at a dog that kept sitting at the dinner table. "Two years at Ironroot Fourth," Master Deng said, releasing Kael's wrist. "The sect's longest stagnation in living memory. Perhaps cultivation simply isn't—" Kael stopped hearing him. Not from anger. That was the strange part. The anger had burned out somewhere around month eight; he knew its shape and limits. This was something else. This was the anger's skeleton, after the flesh was gone. Bare and structural and very, very cold. He felt his qi do something he had no framework for. Ironroot cultivation moved upward—everyone knew this, it was the first lesson, qi rising from the earth through the body's core, building pressure like groundwater finding a fissure. You climbed tiers by sustaining that pressure longer, harder, hotter. …

Head-to-heads

vs Claude Opus 4.8 · vs Grok 4.3 · vs GLM 5.2 · vs Gemini 3.1 Pro · vs DeepSeek V3.2 · vs Mistral Large · vs Llama 4 Maverick

How it works: every pair of models is judged blind on the same scene, with the passages' order flipped so being shown first can't sway it, by GPT-5.4 — a family that isn't on the board, so nobody scores their own side. Each score carries a likely range; overlapping ranges are called a tie. Slop is scored separately by a fixed checklist, not an AI. Full board, every prompt, and the FAQ: thebookfactoryai.com/board. Get each new board by email on the model-drop list.

Writing a book of your own? Book Factory runs the same craft checks on full manuscripts.