TruaceTracing the truth around AITuesday, August 25, 2026
TRV-2026-0883Version 1 · Certified

Written 2026-08-25 14:14:44 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-0883
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-08-25T14:14:44.701183Z
status: published
lens: trace
sector: entertainment
headline: Stylometric comparisons of human versus AI-generated creative writing
dek: This study employs stylometry to investigate whether the creative writing styles of humans and large language models (LLMs) such as GPT-3.5, GPT-4, and Llama 70b can be distinguished through quantitative analysis. A balanced dataset of short stories composed in response to predefined narrative prompts forms the basis of the analysis. Burrows’ Delta, a widely used metric in computational literary studies, is applied to measure stylistic similarity and difference across texts. By focusing on the distribution of th…
gain_title: Quantitative stylometry using Burrows' Delta can reliably separate LLM-generated short stories from human-authored stories, providing a measurable tool for authenticity and authorship checks.
problem_title: LLM-generated creative writing shows higher stylistic uniformity and tight clustering by model, lacking the broader heterogeneity and individual diversity seen in human-authored stories.
trace_subject: stylistic distinctness of LLM versus human short stories generated from predefined narrative prompts
gain_reading: Quantitative stylometry using Burrows' Delta can reliably separate LLM-generated short stories from human-authored stories, providing a measurable tool for authenticity and authorship checks.
gain_evidence: LLM outputs remain statistically and stylistically identifiable as machine-generated.
problem_reading: LLM-generated creative writing shows higher stylistic uniformity and tight clustering by model, lacking the broader heterogeneity and individual diversity seen in human-authored stories.
problem_evidence: LLM outputs, while fluent and coherent, display a higher degree of stylistic uniformity, clustering tightly by model. | Human-authored texts form broader, more heterogeneous clusters, reflecting the diversity of individual expression, writing ability, and interpretive engagement
quick_read: Researchers compared human-authored short stories with stories generated by GPT-3.5, GPT-4, and Llama 70b in response to the same prompts, using Burrows' Delta and clustering methods including hierarchical clustering and multidimensional scaling to visualize stylistic relationships.

The ability to quantitatively identify machine-generated stories matters for debates about authenticity, authorship, and machine creativity in publishing and creative writing, but it remains uncertain how well the signatures hold outside short-form prompted stories and given rare overlaps between GPT-3.5 and human texts.
limitation: Findings are based on a balanced dataset of short stories composed in response to predefined narrative prompts, and occasional overlaps between GPT-3.5 and human texts were observed, limiting generalizability to other genres or open-ended writing.
tag: Dual reading
key_points: Applied Burrows' Delta focusing on distribution of most frequent words to compare latent stylistic fingerprints independent of content. | Used hierarchical clustering and multidimensional scaling on a balanced dataset of short stories written to predefined narrative prompts from humans, GPT-3.5, GPT-4, and Llama 70b. | Found human texts formed broader, more heterogeneous clusters while each LLM clustered tightly by model, with GPT-4 showing greater internal consistency than GPT-3.5.
rundown: The study tested GPT-3.5, GPT-4, and Llama 70b against human authors using the same narrative prompts, applying Burrows' Delta to the most frequent words to isolate style from content.

Clustering results showed GPT-4 with greater internal consistency than GPT-3.5, Llama 70b with similar uniform behavior, and rare overlaps that did not erase the overall separation between human and machine groups.
sources:
- peer_reviewed | Humanities and Social Sciences Communications | https://doi.org/10.1057/s41599-025-05986-3 | 2025-11-11
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
08590170145a3b7f0677617688ca2201fdc0d11f02000fd909cde035f4f86490
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-0883 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.