Fiction Eval leaderboard · model review

Mistral Large for fiction writing — blind-judged review

July 2026: 10 fresh genre scenes, three tries each, judged blind against 7 other frontier models by a different AI family.

Where it lands

#7 of 8 — score 1118 (likely 1054–1167), record 35–162–10 across 207 blind matchups. Slop 17.79 per 1,000 words (5th-cleanest of the field — a separate AI-tell count, not the quality ranking).

#ModelScoreRecordSlop
1Claude Opus 4.81927 (likely 1888–1976)150–14–4314.88
2Claude Sonnet 4.61914 (likely 1874–1961)146–15–4618.57
3Grok 4.31856 (likely 1815–1899)137–28–429.43
4GLM 5.21776 (likely 1732–1822)117–42–4217.53
5Gemini 3.1 Pro1471 (likely 1423–1519)78–104–1313.45
6DeepSeek V3.21251 (likely 1201–1302)52–144–1123.24
7Mistral Large1118 (likely 1054–1167)35–162–1017.79
8Llama 4 Maverick686 (likely 627–741)0–206–122.92

Judged strengths and weaknesses

Traits where Mistral Large scores above the field average: none. Below the field average: voice, pacing, dialogue, genre convention, avoiding AI tells. (Each rated 1–5 by the judge; field average taken across all eight models.)

Sample

LitRPG — “party aftermath” scene; full version on the board.

The screen flickered, a sickly pulse of red light bleeding through the black. My vision swam—no, not my vision. The *system* was resolving, pixel by pixel, like a wound stitching itself shut. **PARTY STATUS: WIPED** **DEATH PENALTY: HARDCORE MODE ENGAGED** The words hung there, jagged and final. I swallowed, but my throat was sand. My fingers twitched toward the log, but I already knew. I’d seen the numbers. Seen the health bars. **Lira (Rogue) – DECEASED** **Garrick (Tank) – DECEASED** **Vex (DPS) – DECEASED** **Kael (Healer) – SURVIVOR** Kael. That was me. The one left standing. The one who’d failed. A new notification blinked into existence, its edges sharp as broken glass. **ACCOUNT FLAGGED: PERMANENT DELETION** **USER: Garrick (Tank) – VIOLATION: TERMINAL DEBT** My breath hitched. *Terminal debt.* That wasn’t just gear loss. That wasn’t just a level drop. That was *gone.* No respawn. No appeal. No second chance. The system log scrolled, merciless: **Garrick has exceeded maximum allowed death penalties.** **Account purged. Character data erased.** **All assets liquidated to …

Head-to-heads

vs Claude Opus 4.8 · vs Claude Sonnet 4.6 · vs Grok 4.3 · vs GLM 5.2 · vs Gemini 3.1 Pro · vs DeepSeek V3.2 · vs Llama 4 Maverick

How it works: every pair of models is judged blind on the same scene, with the passages' order flipped so being shown first can't sway it, by GPT-5.4 — a family that isn't on the board, so nobody scores their own side. Each score carries a likely range; overlapping ranges are called a tie. Slop is scored separately by a fixed checklist, not an AI. Full board, every prompt, and the FAQ: thebookfactoryai.com/board. Get each new board by email on the model-drop list.

Writing a book of your own? Book Factory runs the same craft checks on full manuscripts.