Fiction Eval leaderboard · model review

Claude Opus 4.8 for fiction writing — blind-judged review

July 2026: 10 fresh genre scenes, three tries each, judged blind against 7 other frontier models by a different AI family.

Where it lands

#1 of 8 — score 1927 (likely 1888–1976), record 150–14–43 across 207 blind matchups. Slop 14.88 per 1,000 words (3rd-cleanest of the field — a separate AI-tell count, not the quality ranking).

#ModelScoreRecordSlop
1Claude Opus 4.81927 (likely 1888–1976)150–14–4314.88
2Claude Sonnet 4.61914 (likely 1874–1961)146–15–4618.57
3Grok 4.31856 (likely 1815–1899)137–28–429.43
4GLM 5.21776 (likely 1732–1822)117–42–4217.53
5Gemini 3.1 Pro1471 (likely 1423–1519)78–104–1313.45
6DeepSeek V3.21251 (likely 1201–1302)52–144–1123.24
7Mistral Large1118 (likely 1054–1167)35–162–1017.79
8Llama 4 Maverick686 (likely 627–741)0–206–122.92

Judged strengths and weaknesses

Traits where Claude Opus 4.8 scores above the field average: voice, pacing, dialogue, genre convention, avoiding AI tells. Below the field average: none. (Each rated 1–5 by the judge; field average taken across all eight models.)

Sample

LitRPG — “dungeon-core scene” scene; full version on the board.

I am three days old and I am nearly empty. [MANA: 2/60] The number sits in the dark behind what I am learning to call my sight. I spent the rest — fifty-eight points, everything — hardening the west corridor, packing loose earth into stone because loose earth frightened me. I do not know why it frightened me. I am new. Many things frighten me. Rain has been falling above for a day. I feel it in the soil the way a tooth feels cold water, a long ache down through my single room and my single corridor and into the pebble of crystal that is, apparently, me. Then the ache changes. [PROXIMITY ALERT — LIVING ENTITY] [Entity has entered the Threshold.] Something drags itself in from the rain. I watch — I have no eyes, but the dungeon is my eyes — as it collapses just inside the corridor I hardened. Small. Bent. Ribs like the tines of a broken rake pushing at grey-green skin. It shivers. Water runs …

Head-to-heads

vs Claude Sonnet 4.6 · vs Grok 4.3 · vs GLM 5.2 · vs Gemini 3.1 Pro · vs DeepSeek V3.2 · vs Mistral Large · vs Llama 4 Maverick

How it works: every pair of models is judged blind on the same scene, with the passages' order flipped so being shown first can't sway it, by GPT-5.4 — a family that isn't on the board, so nobody scores their own side. Each score carries a likely range; overlapping ranges are called a tie. Slop is scored separately by a fixed checklist, not an AI. Full board, every prompt, and the FAQ: thebookfactoryai.com/board. Get each new board by email on the model-drop list.

Writing a book of your own? Book Factory runs the same craft checks on full manuscripts.