AI generated novels can look surprisingly different when examined one at a time, but a new analysis suggests their deeper structures become much more alike when dozens are compared together. Researchers generated 80 full-length books using two language models and measured them against 270 human-written novels. The clearest difference was not an obvious giveaway in any single paragraph, but a narrower range of sentence structures across the machine-written collection.
How the AI generated novels were compared
Researchers Mehdy Sedaghat Payam and Justin Quinn describe the experiment in an arXiv preprint on compressed formal variation. They produced 20 novels in a nineteenth-century British realist style using GPT-5.5 Thinking and another 20 in the same style using Qwen3-14B.
They then generated 20 contemporary novels with each model in what the paper calls a zero style. That produced four AI-written groups of 20 books, for a total of 80 novels. The comparison groups comprised 205 human-written nineteenth-century British novels and 65 contemporary human-written novels.
The experiment was not asking whether a detector could identify an artificial paragraph. Instead, it asked whether repeated machine generation could reproduce the breadth of formal variation found across a body of human fiction. A convincing individual imitation and a genuinely diverse literary ecosystem are not the same outcome.
Qwen describes its Qwen3 model family as a capable general-purpose system spanning reasoning and text-generation tasks. The new study does not dispute that such systems can write fluent prose. It examines what happens when that fluency is repeated across many independently generated books.
The strongest signal was sentence structure
The researchers measured average sentence length, readability, punctuation rates, vocabulary-related indicators and a statistical measure of lexical diversity. Across those tests, the most robust result was compression in sentence structure: the generated novels differed less from one another than the human novels did.
That does not mean every sentence had the same number of words or that every synthetic book sounded identical. It means the collection occupied a narrower structural range. The variation between books was more constrained, even where individual titles appeared stylistically distinct.
The paper also reports compressed variation in readability, punctuation and sentence-length variability within novels. Vocabulary measures often pointed in the same direction, although not without exceptions. In particular, one Qwen contemporary-style lexical-diversity measure did not fit the general pattern.
That exception is important because it prevents a simplistic conclusion that every measurable feature converges in every model. Different systems can have distinct average styles while still producing collections with less variation than comparable human-written corpora.
LiveAIWire has previously explored how generative AI is changing the creative industries. This research sharpens the question. The issue is not simply whether a machine can produce a usable book, but what happens to the overall range of books when similar generation systems are used repeatedly.
A convincing page is not the same as a diverse shelf
Imagine evaluating a library by opening a single volume at random. Its paragraphs might sound natural, its punctuation might appear ordinary and its plot might hold together. None of that reveals whether the surrounding shelves contain genuinely varied approaches to rhythm, pacing and sentence construction.
The researchers distinguish this collection-level narrowing from the more familiar search for an AI-writing fingerprint. An individual generated novel may overlap substantially with the stylistic profile of human fiction. The limitation becomes visible when many generated novels are considered as a population.
This distinction matters for publishers, writing platforms and digital retailers. If thousands of individually acceptable books cluster around similar structural habits, readers could experience a subtle flattening of literary variety without being able to identify a single obviously artificial passage.
It also complicates the idea that adding more generated titles necessarily increases cultural diversity. More books can mean more choice in a numerical sense while producing less variation in the formal qualities that distinguish one reading experience from another.
This is not proof of model collapse
The study uses the term variance overclosure to describe a limited range between generated novels. That is not identical to model collapse, the phenomenon in which training repeatedly on synthetic material can degrade future generations of a model.
A separate peer-reviewed Nature study on recursively generated training data examines that distinct problem. The novel experiment does not demonstrate that either model was trained recursively on its own books, or that its capabilities are deteriorating over time.
The authors also distinguish variation across individual measurements from a stronger claim about stable relationships between those measurements. GPT and Qwen displayed different average stylistic profiles, but the study did not find a consistent pattern of cross-measure correlation that would justify treating every aspect of machine-written style as uniformly locked together.
That careful distinction is useful when considering concerns about falling quality in AI-generated content. Structural narrowing is a specific measurable finding. It should not be inflated into an unsupported claim that every synthetic book is unreadable or that every model is collapsing.
What the human comparison can and cannot establish
The human corpora give the research historical and contemporary reference points, but they do not represent all literature. The nineteenth-century group is British, the contemporary category follows the study’s own stylistic definition, and the experiment evaluates only two model families under particular generation conditions.
Public collections such as Project Gutenberg’s volunteer-built digital library illustrate how varied historical writing can be, but digitised archives also reflect editorial, linguistic and preservation choices. No single collection provides a perfect map of human literary possibility.
Prompting, sampling settings, longer-term planning systems and substantial human revision could all affect the results. The paper does not prove that future models cannot generate a broader range, or that an author using AI as one tool in a longer creative process will necessarily produce structurally uniform work.
Nor does measured variation alone determine artistic merit. A deliberately restrained style can be powerful, and two formally similar books can differ profoundly in perspective, character or moral imagination. The research identifies a pattern, not a complete theory of literary value.
Repetition can hide behind convincing surface differences
A machine-generated novel might change its characters, historical setting and plot while preserving a familiar pattern of sentence construction underneath. A casual inspection could therefore exaggerate the amount of creative variety present across a catalogue, particularly when titles, cover art and story premises all appear different.
Human editing could change that outcome, but the amount of intervention matters. Correcting a few awkward phrases is different from substantially reworking rhythm, structure and narrative voice. The experiment does not measure how far an editor would need to go before a generated collection resembles the diversity of a human one.
The result also suggests that publishers should test repeated outputs instead of accepting a provider’s best example. A carefully selected sample can look impressive while disguising how tightly the wider body of generated work clusters around common stylistic habits.
For writers, the more interesting opportunity may be to use AI against its defaults. A person can recognise when a tool repeatedly settles into comfortable rhythms and deliberately push the work towards sharper contrasts, stranger structures or a more distinctive personal voice.
Why creative diversity is a collective question
The practical implication is that creative technology should be evaluated at the level where it will actually operate. A platform planning to generate thousands of books should test whether its catalogue remains varied across repeated production, rather than approving the system after one impressive demonstration.
Publishers could examine distributions of sentence length, readability and stylistic experimentation across many outputs. Writers and editors could deliberately intervene when a model repeatedly defaults to familiar structures. Readers may ultimately value the unusual formal choices that machines smooth away.
As LiveAIWire’s examination of whether machines can display open-ended creativity suggests, creativity is not only about assembling a plausible object. It also concerns the ability to depart from comfortable patterns and sustain meaningful difference.
The study’s strongest message is therefore neither that AI cannot write novels nor that human writing has become obsolete. It is that 80 individually plausible books can still reveal a collective limitation. The deeper question is not whether the machine can fill another shelf, but whether the shelves begin to sound the same.
About the Author
Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.
