Fiction Eval leaderboard · model review

DeepSeek V3.2 for fiction writing — blind-judged review

July 2026: 10 fresh genre scenes, three tries each, judged blind against 7 other frontier models by a different AI family.

Where it lands

#6 of 8 — score 1251 (likely 1201–1302), record 52–144–11 across 207 blind matchups. Slop 23.24 per 1,000 words (8th-cleanest of the field — a separate AI-tell count, not the quality ranking).

#ModelScoreRecordSlop
1Claude Opus 4.81927 (likely 1888–1976)150–14–4314.88
2Claude Sonnet 4.61914 (likely 1874–1961)146–15–4618.57
3Grok 4.31856 (likely 1815–1899)137–28–429.43
4GLM 5.21776 (likely 1732–1822)117–42–4217.53
5Gemini 3.1 Pro1471 (likely 1423–1519)78–104–1313.45
6DeepSeek V3.21251 (likely 1201–1302)52–144–1123.24
7Mistral Large1118 (likely 1054–1167)35–162–1017.79
8Llama 4 Maverick686 (likely 627–741)0–206–122.92

Judged strengths and weaknesses

Traits where DeepSeek V3.2 scores above the field average: none. Below the field average: voice, pacing, dialogue, genre convention, avoiding AI tells. (Each rated 1–5 by the judge; field average taken across all eight models.)

Sample

Cultivation / xianxia — “cultivation beat” scene; full version on the board.

The air in the Grand Ceremony Hall thrummed with spiritual pressure and quiet ambition. Outer disciple Lin Feng stood in his assigned row, his plain grey robes a stark contrast to the shimmering silks of the inner disciples on the raised dais. Today was the annual Ranking Ceremony, where progress was measured, faces were made or lost, and destinies were decided. His gaze was fixed on Inner Disciple Luo Chen, the sect’s rising star. Luo Chen was receiving praise from the Hallmaster for his astonishing leap to the late-stage Qi Condensation realm. “Such pure, vigorous spiritual essence,” the Hallmaster boomed. “A testament to diligent effort and rare talent!” But as Luo Chen bowed, a flare of familiar energy pulsed from his core. It was a unique, slightly discordant resonance Lin Feng had felt a thousand times in his own meridians after his weekly guidance sessions with Mentor Gu. A cold, hollow certainty settled in Lin Feng’s gut. The fatigue that never lifted, the spiritual energy that seemed to leak away …

Head-to-heads

vs Claude Opus 4.8 · vs Claude Sonnet 4.6 · vs Grok 4.3 · vs GLM 5.2 · vs Gemini 3.1 Pro · vs Mistral Large · vs Llama 4 Maverick

How it works: every pair of models is judged blind on the same scene, with the passages' order flipped so being shown first can't sway it, by GPT-5.4 — a family that isn't on the board, so nobody scores their own side. Each score carries a likely range; overlapping ranges are called a tie. Slop is scored separately by a fixed checklist, not an AI. Full board, every prompt, and the FAQ: thebookfactoryai.com/board. Get each new board by email on the model-drop list.

Writing a book of your own? Book Factory runs the same craft checks on full manuscripts.