Fiction Eval leaderboard · model review

Grok 4.3 for fiction writing — blind-judged review

July 2026: 10 fresh genre scenes, three tries each, judged blind against 7 other frontier models by a different AI family.

Where it lands

#3 of 8 — score 1856 (likely 1815–1899), record 137–28–42 across 207 blind matchups. Slop 9.43 per 1,000 words (1st-cleanest of the field — a separate AI-tell count, not the quality ranking).

#ModelScoreRecordSlop
1Claude Opus 4.81927 (likely 1888–1976)150–14–4314.88
2Claude Sonnet 4.61914 (likely 1874–1961)146–15–4618.57
3Grok 4.31856 (likely 1815–1899)137–28–429.43
4GLM 5.21776 (likely 1732–1822)117–42–4217.53
5Gemini 3.1 Pro1471 (likely 1423–1519)78–104–1313.45
6DeepSeek V3.21251 (likely 1201–1302)52–144–1123.24
7Mistral Large1118 (likely 1054–1167)35–162–1017.79
8Llama 4 Maverick686 (likely 627–741)0–206–122.92

Judged strengths and weaknesses

Traits where Grok 4.3 scores above the field average: voice, pacing, dialogue, genre convention, avoiding AI tells. Below the field average: none. (Each rated 1–5 by the judge; field average taken across all eight models.)

Sample

Cozy mystery — “cozy cold open” scene; full version on the board.

The bell over my bookshop door had barely stopped jingling when I crossed the alley to Sweet Haven for my usual pre-festival cinnamon twist. Martha always left the side door unlatched on harvest mornings, the smell of yeast and browned butter guiding half the town inside before the booths were even set up. Today the kitchen was quiet except for the low hum of the walk-in. “Martha?” I called, wiping my hands on my cardigan. No answer, just the distant clang of the festival crew raising the bunting outside. I pushed the heavy door wider and found her on the floor between the racks of cooling pumpkin loaves, apron still tied, cheeks pink as if she’d only just stepped in to fetch butter. Doc Ellis arrived with the constable, both of them moving carefully around the trays so as not to disturb the order Martha had kept for thirty years. They spoke in the low, respectful tones people use when the deceased is beloved. Someone mentioned the faulty latch again, …

Head-to-heads

vs Claude Opus 4.8 · vs Claude Sonnet 4.6 · vs GLM 5.2 · vs Gemini 3.1 Pro · vs DeepSeek V3.2 · vs Mistral Large · vs Llama 4 Maverick

How it works: every pair of models is judged blind on the same scene, with the passages' order flipped so being shown first can't sway it, by GPT-5.4 — a family that isn't on the board, so nobody scores their own side. Each score carries a likely range; overlapping ranges are called a tie. Slop is scored separately by a fixed checklist, not an AI. Full board, every prompt, and the FAQ: thebookfactoryai.com/board. Get each new board by email on the model-drop list.

Writing a book of your own? Book Factory runs the same craft checks on full manuscripts.