Fiction Eval leaderboard · model review

Llama 4 Maverick for fiction writing — blind-judged review

July 2026: 10 fresh genre scenes, three tries each, judged blind against 7 other frontier models by a different AI family.

Where it lands

#8 of 8 — score 686 (likely 627–741), record 0–206–1 across 207 blind matchups. Slop 22.92 per 1,000 words (7th-cleanest of the field — a separate AI-tell count, not the quality ranking).

#ModelScoreRecordSlop
1Claude Opus 4.81927 (likely 1888–1976)150–14–4314.88
2Claude Sonnet 4.61914 (likely 1874–1961)146–15–4618.57
3Grok 4.31856 (likely 1815–1899)137–28–429.43
4GLM 5.21776 (likely 1732–1822)117–42–4217.53
5Gemini 3.1 Pro1471 (likely 1423–1519)78–104–1313.45
6DeepSeek V3.21251 (likely 1201–1302)52–144–1123.24
7Mistral Large1118 (likely 1054–1167)35–162–1017.79
8Llama 4 Maverick686 (likely 627–741)0–206–122.92

Judged strengths and weaknesses

Traits where Llama 4 Maverick scores above the field average: none. Below the field average: voice, pacing, dialogue, genre convention, avoiding AI tells. (Each rated 1–5 by the judge; field average taken across all eight models.)

Sample

Science fiction — “quiet first contact” scene; full version on the board.

Elara hunched over the console, fingers flying across the keyboard as she scrubbed through the latest batch of data from the ATA-12 receiver. The screen flickered with a waterfall display of radio frequency interference, a cacophony of human noise and natural static that she'd spent countless hours cataloging. It was a thankless task, but someone had to do it – the SETI project's credibility hinged on proving that the signal they were searching for wasn't just some mundane terrestrial interference. As she worked, her mind wandered to the puzzle that had been nagging her for weeks. The "noise" – a series of brief, narrowband pulses at 1420 MHz – had been showing up in the data with maddening consistency. She'd run it through the signal processing algorithms, and while it didn't quite fit the profile of a technosignature, it was...different. Elara's thoughts drifted to the implications: a transmission from a distant civilization, perhaps? The idea sent a shiver down her spine. Her train of thought was interrupted by the console …

Head-to-heads

vs Claude Opus 4.8 · vs Claude Sonnet 4.6 · vs Grok 4.3 · vs GLM 5.2 · vs Gemini 3.1 Pro · vs DeepSeek V3.2 · vs Mistral Large

How it works: every pair of models is judged blind on the same scene, with the passages' order flipped so being shown first can't sway it, by GPT-5.4 — a family that isn't on the board, so nobody scores their own side. Each score carries a likely range; overlapping ranges are called a tie. Slop is scored separately by a fixed checklist, not an AI. Full board, every prompt, and the FAQ: thebookfactoryai.com/board. Get each new board by email on the model-drop list.

Writing a book of your own? Book Factory runs the same craft checks on full manuscripts.