Skip to content

Build1 publisher3 min readPublished

The AI book problem is arithmetic: catalog up 38x, revenue up 9x

An analysis of 14,419 self-published Amazon e-books found revenue per title falling even where no AI text was detected. That points at flooding, not reader preference.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying The AI book problem is arithmetic: catalog up 38x, revenue up 9x
Generated illustration

What happened

  • Researchers analyzed 14,419 randomly selected self-published e-books released between January 2023 and March 2026.
  • The analysis concludes that AI-generated titles are displacing human authors through sheer volume, not quality.
  • Revenue per book is falling even for titles where no AI text was detected.
  • Books were sorted into three bands by share of text flagged as AI-generated: none, light (up to 25 percent), and substantial (over 25 percent).
  • Books with substantial AI content make up 20 percent of the catalog studied but account for only 12.1 percent of sales and 11.3 percent of revenue.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

An analysis of 14,419 randomly selected self-published Amazon e-books released between January 2023 and March 2026 concludes that AI-generated titles are displacing human authors through volume rather than by winning readers [1] [2]. The load-bearing finding is that revenue per book fell even for titles where the detector found no AI text at all [3], which shifts the diagnosis from quality to congestion.

The quality story holds up on its own terms, and it is the less interesting one. Books with substantial AI content, meaning more than 25 percent of the text flagged, make up 20 percent of the catalog but only 12.1 percent of sales and 11.3 percent of revenue [4] [5]. Books with no detected AI text are 62.9 percent of the catalog and 72.5 percent of revenue [6]. Per unit of shelf space, in other words, a substantially-AI title earns about 0.57 times its proportional share of revenue and a clean title about 1.15 times [1]. Read alone, that says slop stays at the bottom [7].

Now the arithmetic underneath it. Between Q1 2023 and Q1 2026 the cumulative catalog grew 38.3x, the number of titles selling per quarter grew 19.2x, and quarterly revenue grew 8.9x [8]. Revenue per catalogued title therefore fell to roughly 23 percent of its starting level, and revenue per title that actually sold in a given quarter fell to roughly 46 percent [2] [3]. Everyone is fishing in a pool that grew nine times while the boats grew nineteen to thirty-eight times.

The per-book decline is not an averaging artifact. Comparing 2023 and 2025 releases over the same post-release window, revenue per book dropped in six of eight genres; restricted to books with no detected AI text, it dropped in seven of eight [9] [10]. The exception is Fantasy/Supernatural/Horror, where AI text arrived latest and gained the least traction, and where revenue per clean title rose 35 percent [11]. The researchers argue that reversal weighs against a general market downturn as the explanation [12], and they label the effect dilution while stating plainly that their comparisons are observational and associational, not experimental proof of causation [13].

Supply is concentrated. Of 385 author identities that published more substantial-AI titles after their first, 287, or about 75 percent, increased their monthly output afterward [14] [4]. The highest-grossing pseudonym took $1.7 million in gross revenue before platform fees across eight titles, an average of about $213,000 a title [15] [5]. The single highest-grossing substantial-AI book earned $643,000 on 80,431 copies, roughly $8 a copy [16] [6]. And the top of the chart is no longer insulated: new Top 25 entries with substantial AI content went from near zero to 31 percent, while quarter-to-quarter persistence for clean books in the Top 25 fell as low as about 28 percent before settling near 62 percent [17] [18].

Two caveats to hold onto. Classification used full text with Pangram v3.3, whose developers report a 0.04 percent false-positive rate, and Pangram 4 has since shipped [19] [20] [21]. Sales came from an internal dataset held by one of the five major US publishers, tracking about 500,000 Amazon titles and roughly 95 percent of daily e-book sales, according to the researchers [22].

Watch the Kindle Unlimited split: in high-KU genres the revenue-share lead of clean books is 8.4 percentage points smaller, though the authors attribute that to genre traits and decline to blame the subscription pool itself [23] [24]. Watch whether the dilution result survives reclassification under Pangram 4 [21]. And watch the provenance work the write-up leaves unfinished, an overlap test built on the Allen Institute's infini-gram and Google Books whose results the source cuts off mid-sentence [25].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories