← Research

LLM Creative Writing Benchmark

Original study 2025 · longitudinal re-run June 2026 · ongoing

When you ask a model for the same story ten times, how differently does it write? The benchmark sends one science-fiction prompt to each model ten times (1,500-word target) and measures verbatim similarity, semantic similarity, vocabulary, named entities, and text structure. The 2025 study covered GPT-4o, o1, Gemini 2.0/2.5, and Claude 3.5/3.7; the 2026 re-run covers Gemini 3 Flash and Pro, Claude Opus 4.8, Sonnet, and Fable 5, and GPT-5.5.

Headline findings (June 2026 run)

Method and caveats

The 2026 runs were generated through local agent CLIs (claude, gemini, codex) in headless mode, not the provider APIs used in 2025. The CLIs do not forward sampling parameters, so the 2025 temperature of 0.7 no longer applies and each CLI used its own default sampling. The 2026 figures are therefore same-prompt, same-model-family comparisons rather than strictly controlled replications; the natural next step is an API re-run at a controlled temperature.

This work feeds directly into the narrative engines behind DungeonGPT and our writing tools. It also sets up a natural follow-up question: how do these numbers compare against fiction written by humans? We built a human baseline for exactly that. See Human Masterworks vs AI Fiction.