← Research

Human Masterworks vs AI Fiction

Masters corpus · 26 books scored · July 2026 · ongoing

Our creative writing benchmark measures how much language models vary when they retell the same story, from each other and from themselves. It answers "how different are the models." It does not answer the question underneath it: how does any of this compare to real, acknowledged good writing?

Scoring machine prose is hard without a yardstick. A number like "lexical diversity 91" or "cliche density 0.05 per thousand words" only means something once you know the range that skilled, published fiction actually occupies. So we scored a corpus of 26 canonical, public-domain masterworks (Austen, Dickens, Conrad, Collins, Eliot, Tolstoy, Dumas, Stoker, Wells, Haggard, Buchan, Sabatini, and others) on the same tooling we point at model output. The result is a reference band: the values real literature falls in, so AI-generated text can be measured against books rather than against a guess about what "good" should look like.

Two benchmarks

Headline findings

Method and validation

Every book was extracted from raw Project Gutenberg text to a verified, canonical Markdown file and then frozen, so a later extractor change cannot silently move already-published scores. The narrative-dynamics judge (DeepSeek) was validated against a second model (Claude Haiku 4.5) on Dracula's 28-chapter tension curve: the two judges agree at Pearson r = 0.86, differing by less than a point on average on the 0 to 10 scale. Every claim links back to the source texts and the analysis code, the same standard as the rest of our research.

Honest caveats