stylistic distinctness of LLM versus human short stories generated from predefined narrative prompts
Source article: Stylometric comparisons of human versus AI-generated creative writing
Abstract: This study employs stylometry to investigate whether the creative writing styles of humans and large language models (LLMs) such as GPT-3.5, GPT-4, and Llama 70b can be distinguished through quantitative analysis. A balanced dataset of short stories composed in response to predefined narrative prompts forms the basis of the analysis. Burrows’ Delta, a widely used metric in computational literary studies, is applied to measure stylistic similarity and difference across texts. By focusing on the distribution of th…
Negative state: both sides are scored from claims and sources, not community votes.
Mr. Montague Chambers' address to the jury, in the case of Gardner versus Godfrey, from the short-hand notes of J. Freeman by Freeman, J Chambers, Montague Royal College of Physicians of Edinburgh. Public domain
Researchers compared human-authored short stories with stories generated by GPT-3.5, GPT-4, and Llama 70b in response to the same prompts, using Burrows' Delta and clustering methods including hierarchical clustering and multidimensional scaling to visualize stylistic relationships.
The ability to quantitatively identify machine-generated stories matters for debates about authenticity, authorship, and machine creativity in publishing and creative writing, but it remains uncertain how well the signatures hold outside short-form prompted stories and given rare overlaps between GPT-3.5 and human texts.
- Applied Burrows' Delta focusing on distribution of most frequent words to compare latent stylistic fingerprints independent of content.
- Used hierarchical clustering and multidimensional scaling on a balanced dataset of short stories written to predefined narrative prompts from humans, GPT-3.5, GPT-4, and Llama 70b.
- Found human texts formed broader, more heterogeneous clusters while each LLM clustered tightly by model, with GPT-4 showing greater internal consistency than GPT-3.5.
Quantitative stylometry using Burrows' Delta can reliably separate LLM-generated short stories from human-authored stories, providing a measurable tool for authenticity and authorship checks.
LLM-generated creative writing shows higher stylistic uniformity and tight clustering by model, lacking the broader heterogeneity and individual diversity seen in human-authored stories.
The rundown
The study tested GPT-3.5, GPT-4, and Llama 70b against human authors using the same narrative prompts, applying Burrows' Delta to the most frequent words to isolate style from content.
Clustering results showed GPT-4 with greater internal consistency than GPT-3.5, Llama 70b with similar uniform behavior, and rare overlaps that did not erase the overall separation between human and machine groups.
Findings are based on a balanced dataset of short stories composed in response to predefined narrative prompts, and occasional overlaps between GPT-3.5 and human texts were observed, limiting generalizability to other genres or open-ended writing.
Sources
- Peer-reviewedHumanities and Social Sciences Communications2025-11-11
How should this claim be treated?
ace
The debate