Build1 distinct publisher3 min readUpdated
An analysis of 14,419 self-published Amazon e-books found revenue per title falling even where no AI text was detected. That points at flooding, not reader preference.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
An analysis of 14,419 randomly selected self-published Amazon e-books released between January 2023 and March 2026 concludes that AI-generated titles are displacing human authors through volume rather than by winning readers [1] [2]. The load-bearing finding is that revenue per book fell even for titles where the detector found no AI text at all [3], which shifts the diagnosis from quality to congestion.
The quality story holds up on its own terms, and it is the less interesting one. Books with substantial AI content, meaning more than 25 percent of the text flagged, make up 20 percent of the catalog but only 12.1 percent of sales and 11.3 percent of revenue [4] [5]. Books with no detected AI text are 62.9 percent of the catalog and 72.5 percent of revenue [6]. Per unit of shelf space, in other words, a substantially-AI title earns about 0.57 times its proportional share of revenue and a clean title about 1.15 times [1]. Read alone, that says slop stays at the bottom [7].
Now the arithmetic underneath it. Between Q1 2023 and Q1 2026 the cumulative catalog grew 38.3x, the number of titles selling per quarter grew 19.2x, and quarterly revenue grew 8.9x [8]. Revenue per catalogued title therefore fell to roughly 23 percent of its starting level, and revenue per title that actually sold in a given quarter fell to roughly 46 percent [2] [3]. Everyone is fishing in a pool that grew nine times while the boats grew nineteen to thirty-eight times.
The per-book decline is not an averaging artifact. Comparing 2023 and 2025 releases over the same post-release window, revenue per book dropped in six of eight genres; restricted to books with no detected AI text, it dropped in seven of eight [9] [10]. The exception is Fantasy/Supernatural/Horror, where AI text arrived latest and gained the least traction, and where revenue per clean title rose 35 percent [11]. The researchers argue that reversal weighs against a general market downturn as the explanation [12], and they label the effect dilution while stating plainly that their comparisons are observational and associational, not experimental proof of causation [13].
Supply is concentrated. Of 385 author identities that published more substantial-AI titles after their first, 287, or about 75 percent, increased their monthly output afterward [14] [4]. The highest-grossing pseudonym took $1.7 million in gross revenue before platform fees across eight titles, an average of about $213,000 a title [15] [5]. The single highest-grossing substantial-AI book earned $643,000 on 80,431 copies, roughly $8 a copy [16] [6]. And the top of the chart is no longer insulated: new Top 25 entries with substantial AI content went from near zero to 31 percent, while quarter-to-quarter persistence for clean books in the Top 25 fell as low as about 28 percent before settling near 62 percent [17] [18].
Two caveats to hold onto. Classification used full text with Pangram v3.3, whose developers report a 0.04 percent false-positive rate, and Pangram 4 has since shipped [19] [20] [21]. Sales came from an internal dataset held by one of the five major US publishers, tracking about 500,000 Amazon titles and roughly 95 percent of daily e-book sales, according to the researchers [22].
Watch the Kindle Unlimited split: in high-KU genres the revenue-share lead of clean books is 8.4 percentage points smaller, though the authors attribute that to genre traits and decline to blame the subscription pool itself [23] [24]. Watch whether the dilution result survives reclassification under Pangram 4 [21]. And watch the provenance work the write-up leaves unfinished, an overlap test built on the Allen Institute's infini-gram and Google Books whose results the source cuts off mid-sentence [25].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Books were sorted into three bands by share of text flagged as AI-generated: none, light (up to 25 percent), and substantial (over 25 percent).
To measure how much successful AI books overlap with rare language from existing works, the researchers used the Allen Institute for AI's infini-gram tool together with Google Books; the source text ends mid-sentence at this point without reporting the result.
Unlike earlier studies that tried to detect AI text from short book previews, the researchers classified each book based on its full text.
They used the Pangram v3.3 detector, whose developers report a false-positive rate of 0.04 percent.
Pangram 4 has since been released.
Daily sales figures came from an internal dataset maintained by one of the five major US publishers, which tracks about 500,000 Amazon titles and covers roughly 95 percent of all e-books sold daily on the platform, according to the researchers.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Large quantified dataset, one publisher's account, self-limited causality
The underlying work is unusually concrete for this beat: 14,419 randomly selected titles, full-text rather than preview classification, daily sales from a dataset the researchers say covers ~95 percent of Amazon e-book sales, and internally consistent multiples (38.3x / 19.2x / 8.9x) plus genre-level and Top 25 breakdowns. It is discounted because everything reaches us through one publisher with no study link or peer-review status, detector accuracy is vendor-reported on a now-superseded version, and the authors themselves disclaim causal inference.
One in five studied titles substantially AI-written; 31 percent of new Top 25 entries
Adoption of generative text in self-publishing is directly measured, not projected: 20 percent of the studied catalog carries substantial AI content, substantial-AI titles now supply 31 percent of new Top 25 entries, and identified prolific identities raised output after adopting AI, with individual titles grossing in the hundreds of thousands. It is not higher because the measurement is one catalog snapshot on one platform, detected by a single tool.
Causal 'flooding and tanking' framing outruns associational findings
The measured pattern is real and quantified, but the presentation is stronger than the study's own claim: the headline asserts AI books are tanking human authors' sales while the authors explicitly describe observational, associational comparisons. Conversely, the article does volunteer the counterweight that substantial-AI titles underperform on revenue share, and the Fantasy exception is reported rather than buried, so the gap is moderate rather than severe.
Publisher-held sales data, vendor-reported detector accuracy, live copyright litigation
Several interested parties sit inside the evidence chain: the daily sales dataset is maintained by one of the five major US publishers, which is a competitor to self-publishing and a litigant class in AI copyright disputes; detector accuracy is reported by the detector's developers; and the article states the findings feed into ongoing copyright cases, with a quoted researcher framing the rare-language method as helping debunk AI companies' training-versus-reading defense. None of these invalidate the numbers, but they shape which numbers exist and how they are framed.
Consistent single-publisher account awaiting independent corroboration
Internal consistency is good and the figures are specific, so the headline direction — far more titles chasing slower-growing revenue — is credible. Confidence stays mid-range because the cluster has exactly one publisher, no direct study reference, an unverified detector and dataset, an unreconciled causal framing, and a source body that breaks off mid-section on copyright implications.
science
Text watermarks land on 2 December. The detection they imply does not.1 distinct publisher
invest
Behind-the-meter gas is the data center buildout's real cost: 318 Mt a year1 distinct publisher
build
PerceptionBench puts a number on the step your pipeline treats as free1 distinct publisher
leadership
Sponsorship Is Not Retention: What An Amazon H-1B Win Still Does Not Buy1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 15, 2026