TruaceTracing the truth around AITuesday, August 25, 2026
Media & Arts·The Trace·Dual reading·Published 2026-08-25

stylistic distinctness of LLM versus human short stories generated from predefined narrative prompts

Source article: Stylometric comparisons of human versus AI-generated creative writing

Abstract: This study employs stylometry to investigate whether the creative writing styles of humans and large language models (LLMs) such as GPT-3.5, GPT-4, and Llama 70b can be distinguished through quantitative analysis. A balanced dataset of short stories composed in response to predefined narrative prompts forms the basis of the analysis. Burrows’ Delta, a widely used metric in computational literary studies, is applied to measure stylistic similarity and difference across texts. By focusing on the distribution of th…

TRV-2026-0883Peer-reviewedPermanent record — cite & verify
Trace impact reading

Negative state: both sides are scored from claims and sources, not community votes.

P 68The P score combines the specificity and measured human impact of the grounded problem claim with the strength of this Trace’s cited sources.G 63The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.
Stylometric comparisons of human versus AI-generated creative writing

Mr. Montague Chambers' address to the jury, in the case of Gardner versus Godfrey, from the short-hand notes of J. Freeman by Freeman, J Chambers, Montague Royal College of Physicians of Edinburgh. Public domain

The quick read

Researchers compared human-authored short stories with stories generated by GPT-3.5, GPT-4, and Llama 70b in response to the same prompts, using Burrows' Delta and clustering methods including hierarchical clustering and multidimensional scaling to visualize stylistic relationships.

The ability to quantitatively identify machine-generated stories matters for debates about authenticity, authorship, and machine creativity in publishing and creative writing, but it remains uncertain how well the signatures hold outside short-form prompted stories and given rare overlaps between GPT-3.5 and human texts.

Main points
  • Applied Burrows' Delta focusing on distribution of most frequent words to compare latent stylistic fingerprints independent of content.
  • Used hierarchical clustering and multidimensional scaling on a balanced dataset of short stories written to predefined narrative prompts from humans, GPT-3.5, GPT-4, and Llama 70b.
  • Found human texts formed broader, more heterogeneous clusters while each LLM clustered tightly by model, with GPT-4 showing greater internal consistency than GPT-3.5.
Gain

Quantitative stylometry using Burrows' Delta can reliably separate LLM-generated short stories from human-authored stories, providing a measurable tool for authenticity and authorship checks.

Problem

LLM-generated creative writing shows higher stylistic uniformity and tight clustering by model, lacking the broader heterogeneity and individual diversity seen in human-authored stories.

The rundown

The study tested GPT-3.5, GPT-4, and Llama 70b against human authors using the same narrative prompts, applying Burrows' Delta to the most frequent words to isolate style from content.

Clustering results showed GPT-4 with greater internal consistency than GPT-3.5, Llama 70b with similar uniform behavior, and rare overlaps that did not erase the overall separation between human and machine groups.

What this doesn’t fix

Findings are based on a balanced dataset of short stories composed in response to predefined narrative prompts, and occasional overlaps between GPT-3.5 and human texts were observed, limiting generalizability to other genres or open-ended writing.

Sources

Reader signal

How should this claim be treated?

The debate