Fiction Eval leaderboard · model review

Gemini 3.1 Pro for fiction writing — blind-judged review

July 2026: 10 fresh genre scenes, three tries each, judged blind against 7 other frontier models by a different AI family.

Where it lands

#5 of 8 — score 1471 (likely 1423–1519), record 78–104–13 across 195 blind matchups. Slop 13.45 per 1,000 words (2nd-cleanest of the field — a separate AI-tell count, not the quality ranking).

#ModelScoreRecordSlop
1Claude Opus 4.81927 (likely 1888–1976)150–14–4314.88
2Claude Sonnet 4.61914 (likely 1874–1961)146–15–4618.57
3Grok 4.31856 (likely 1815–1899)137–28–429.43
4GLM 5.21776 (likely 1732–1822)117–42–4217.53
5Gemini 3.1 Pro1471 (likely 1423–1519)78–104–1313.45
6DeepSeek V3.21251 (likely 1201–1302)52–144–1123.24
7Mistral Large1118 (likely 1054–1167)35–162–1017.79
8Llama 4 Maverick686 (likely 627–741)0–206–122.92

Judged strengths and weaknesses

Traits where Gemini 3.1 Pro scores above the field average: dialogue, genre convention. Below the field average: voice, pacing, avoiding AI tells. (Each rated 1–5 by the judge; field average taken across all eight models.)

Sample

Thriller — “thriller cold open” scene; full version on the board.

The graveyard shift at the Seattle Air Route Traffic Control Center was supposed to be sixty percent caffeine and forty percent boredom. At 0314 hours, Martin’s sector was a black void on the En Route Automation Modernization display, save for a single FedEx heavy crawling down the coast toward Oakland. Then the ERAM screen refreshed, and a new data block bloomed in Sector 4. Martin rested his hand on the trackball. He didn’t spill his coffee. He didn’t gasp. Twenty years of staring at radar beat the panic out of you. He simply rolled the cursor over the primary return and clicked to interrogate the track. No flight plan was attached. Unscheduled. He keyed his transmitter. "Unknown aircraft, Seattle Center. You are in controlled airspace. Squawk ident and state intentions." Silence. The radar swept again. The transponder data populated the block. Martin frowned, leaning closer to the high-definition glass. `ALT FL600` `SPD 0950` `SQK 8888` Martin blinked. He rubbed his eyes and looked again. Flight Level 600 was sixty thousand …

Head-to-heads

vs Claude Opus 4.8 · vs Claude Sonnet 4.6 · vs Grok 4.3 · vs GLM 5.2 · vs DeepSeek V3.2 · vs Mistral Large · vs Llama 4 Maverick

How it works: every pair of models is judged blind on the same scene, with the passages' order flipped so being shown first can't sway it, by GPT-5.4 — a family that isn't on the board, so nobody scores their own side. Each score carries a likely range; overlapping ranges are called a tie. Slop is scored separately by a fixed checklist, not an AI. Full board, every prompt, and the FAQ: thebookfactoryai.com/board. Get each new board by email on the model-drop list.

Writing a book of your own? Book Factory runs the same craft checks on full manuscripts.