LLM Creative Story-Writing Benchmark

September 5, 2026 · View on GitHub

This benchmark compares short stories written to the same constrained creative briefs. Separate evaluator models read matched story pairs and choose which one is better. Those choices are combined into a relative comparison score.

Higher scores mean stronger performance against the other models tested. Scores are relative, not grades: zero is near the middle of this comparison set, and overlapping uncertainty ranges can indicate similarly rated models.


Current Results

Comparison ratings

Leaderboard

Current comparison set:

  • 50 rated models
  • 773 direct model pairings
  • 79,507 evaluator judgments
  • the rating combines compatible evaluator-v2 and evaluator-v3 evidence after bridge validation
  • the chart focuses on selected current models; the table retains all rated models for historical comparison
  • striped bars and markers identify models that completed fewer than 400 stories
RankModelComparison scoreEstimated win chanceUncertainty range
1Claude Opus 5 (xhigh)4.095%3.9 to 4.1
2Claude Fable 5.1 (high)4.094%3.9 to 4.1
3GPT-6 Astra (high)3.592%3.4 to 3.6
4GLM-5.3 (max)3.189%3.0 to 3.3
5Claude Fable 5 (high)§3.188%3.0 to 3.1
6Kimi K32.785%2.6 to 2.8
7GPT-5.6 Sol (xhigh)2.784%2.6 to 2.7
8GPT-5.5 (xhigh)2.684%2.6 to 2.7
9GPT-5.6 Sol (high)2.583%2.4 to 2.6
10GPT-5.4 (medium)2.279%2.1 to 2.3
11GPT-5.4 (xhigh)2.279%2.1 to 2.3
12Claude Opus 4.7 (high)†2.077%1.9 to 2.1
13Claude Sonnet 4.6 Thinking 16K1.976%1.8 to 2.0
14Claude Opus 4.6 Thinking 16K1.469%1.2 to 1.5
15Claude Opus 4.8 (xhigh)1.064%0.9 to 1.1
16Muse Spark 1.1 (high)1.064%0.9 to 1.1
17Muse Spark 1.3 (high)0.861%0.7 to 0.9
18DeepSeek V4 Pro (high)0.761%0.6 to 0.9
19GLM-5.2 (max)0.660%0.5 to 0.7
20GPT-5.2 (medium)0.558%0.4 to 0.7
21Qwen 3.8 Max¶0.557%0.3 to 0.6
22Claude Opus 4.8 (high)‡0.557%0.3 to 0.6
23Kimi K2.60.355%0.2 to 0.4
24Muse Spark 1.2 (high)0.252%0.0 to 0.3
25MiniMax-M30.152%0.0 to 0.3
26Mistral Medium 3.1-0.247%-0.3 to 0.0
27DeepSeek V4 Pro Preview-0.345%-0.4 to -0.2
28Xiaomi MiMo V2.5 Pro-0.444%-0.6 to -0.3
29Qwen 3 Max Preview-0.543%-0.6 to -0.3
30Qwen3.8-27B^-0.543%-0.7 to -0.3
31Gemini 3.7 Flash (high)-0.542%-0.7 to -0.4
32Qwen 3.6 Max Preview-0.739%-0.9 to -0.6
33GLM-5.1-0.838%-1.0 to -0.7
34Kimi K2.5 Thinking-0.937%-1.0 to -0.7
35Xiaomi MiMo V2 Pro-1.035%-1.2 to -0.8
36Baidu Ernie 5.1-1.035%-1.2 to -0.9
37Mistral Large 3-1.627%-1.7 to -1.5
38Gemma 4 31B Reasoning-1.726%-1.8 to -1.6
39Gemini 3.5 Flash-1.824%-1.9 to -1.7
40ByteDance Seed2.0 Pro-1.824%-2.0 to -1.7
41Gemini 3.1 Pro Preview-2.122%-2.1 to -1.9
42Qwen 3.6 Plus-2.121%-2.3 to -1.9
43Mistral Medium 3.5-2.418%-2.5 to -2.2
44Qwen 3.7 Max-2.517%-2.6 to -2.3
45DeepSeek V3.2-2.715%-3.0 to -2.4
46Grok 4.6 (high)-2.714%-2.9 to -2.6
47GPT-OSS-120B-3.012%-3.1 to -2.8
48MiniMax-M2.7-3.68%-3.7 to -3.4
49Grok 4.3-4.15%-4.3 to -3.9
50Grok 4.5 (high)-4.92%-5.0 to -4.8

Coverage Note

  • † Claude Opus 4.7 completed 347 of 400 stories. Only completed stories were compared.
  • ‡ Claude Opus 4.8 high completed 399 of 400 stories. Only completed stories were compared.
  • § Claude Fable 5 high completed 395 of 400 stories. Only completed stories were compared.
  • ¶ Qwen 3.8 Max completed 398 of 400 stories after one audited recovery pass. Only completed stories were compared.
  • ^ Qwen3.8-27B completed 389 of 400 stories after one audited recovery pass. Only completed stories were compared.

Head-to-Head Comparisons

Pairwise margin heatmap

Read each cell by row. Red means the row model performed better, blue means the column model performed better, and grey means the models were not directly compared. Near-white cells indicate close results. Both axes follow the leaderboard order.


Diagnostics

Evaluator Agreement

Evaluator agreement matrix

This chart shows how similarly the evaluator models scored the same story pairs. Values closer to 1 indicate stronger agreement; 0 means no consistent relationship, and negative values mean opposing scoring patterns. Its colorblind-safe scale uses orange for negative relationships and blue for positive relationships.

Word Count Compliance

Word count compliance

Each dot is one story, and each diamond is a model's average length. Thin vertical lines show uncertainty around the averages. The shaded band is the 600-800-word target. This chart measures story length, not writing quality.


What Is Measured

Every story must meaningfully incorporate ten required elements:

  • character
  • object
  • concept
  • attribute
  • action
  • method
  • setting
  • timeframe
  • motivation
  • tone

Candidate combinations are proposed for coherence and originality, then independently rated; the strongest set for each seed becomes a fixed brief used by every writer model. A typical brief might combine a neutron-star researcher, butterfly-wing dust, gradual change, a storm-damaged greenhouse, "after the flood," and kindled humility.

Evaluators reward integration rather than keyword inclusion: the required object should affect the plot, the motivation should produce a consequential choice, and the tone should shape the story's development. Both stories in every comparison answer the same brief, holding prompt difficulty constant. Evaluators also consider prose, coherence, character, originality, and overall effectiveness. The public score combines their choices; it is not an average 0-10 grade.

Because the combinations are pre-screened for creative potential, the benchmark measures story construction under deliberately combinable constraints—not completely free-form writing or recovery from arbitrary incoherent prompts.


Method Summary

  1. Generate stories in the benchmark format.
  2. Build matched story-comparison prompts for models that wrote to the same required elements.
  3. Show each pair in both story orders to reduce first- or second-position effects.
  4. Repeat comparisons across evaluators and combine their choices.
  5. When an evaluator roster changes, validate shared prompts and bridge matchups before combining evidence.
  6. Calculate relative model scores and their uncertainty ranges.

Qualitative Pair Reports

New Models Compared with Their Predecessors collects nine release-to-predecessor reports on a separate page.


Public Artifacts

The published bundle includes the story prompts and generated story text files for models visible in the public comparison charts, plus the linked qualitative reports. Prompt files are under prompts_wc/; model outputs are under stories_wc/<model>/.

Public benchmark data provides machine-readable leaderboard, head-to-head, pair-story, and evaluator-diagnostic tables. It also links to an immutable data release containing the exact evaluator prose for both story orders, excluded or superseded responses, and the referenced story and prompt texts. Provider envelopes, request identifiers, internal paths, and implementation code are not published.


Archived Absolute Ratings

Earlier versions of this benchmark used absolute 0-10 rubric ratings rather than direct story comparisons. Those results remain historical context, but the current public quality ranking should use the pairwise comparison results above.


Recent Updates

  • September 5, 2026: Added GPT-6 Astra high, Muse Spark 1.3.
  • September 2, 2026: Added Claude Fable 5.1.
  • August 23, 2026: Published machine-readable comparison data and an auditable evaluator-prose release.
  • August 20, 2026: Added Muse Spark 1.2 high, DeepSeek V4 Pro high, Qwen 3.8 Max, Gemini 3.7 Flash high, and Grok 4.6 high. Added predecessor reports.
  • July 25, 2026: Added Claude Opus 5.
  • July 18, 2026: Added Kimi K3, updated evaluators.
  • July 14, 2026: Added GPT-5.6, Muse Spark 1.1 high, and Grok 4.5.
  • July 9, 2026: Added Grok 4.5.
  • June 9, 2026: Added Claude Fable 5.
  • May 29, 2026: Added Claude Opus 4.8 high and xhigh.
  • May 26, 2026: Ernie 5.1, Qwen 3.7 Max, Mistral Medium 3.5, and Grok 4.3 added.
  • May 20, 2026: Added Gemini 3.5 Flash.
  • Apr 29, 2026: Refreshed the leaderboard with newer models, including GPT-5.5, Kimi K2.6, DeepSeek V4 Pro, Xiaomi MiMo V2.5 Pro, Qwen 3 Max Preview, Gemini 3.1 Pro Preview, ByteDance Seed2.0 Pro, Qwen 3.6 Max Preview, and MiniMax-M2.7.

Multi-agent benchmarks:

Other benchmarks:

Follow @lechmazur on X for other benchmarks and updates.