Build1 publisher2 min readPublished
Wild AI text, about 31% of web tokens, gets a harm term in a new scaling law
Pangram Labs and UMass Amherst researchers put AI-generated text at about 31% of August 2026 web tokens and fit a scaling law for when it starts to hurt. For teams that pretrain on web crawls, AI text becomes a measured line item in the data budget, with a curve for its marginal value.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The Pangram detector's count of AI-generated web text rose from 27.5% in June 2026 to over 31% two months later.
- The UMass Amherst team pretrained 800 separate language models, varying the ratio of human to AI-generated tokens in each.
- Once a model had a large human-text budget, adding AI tokens past a threshold raised its loss on human text instead of lowering it.
- The researchers released WildAI, an 83-billion-token corpus labeled by topic, format, and whether each source is human or AI.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision One AI-text cutoff applied to both a data-starved run and a run with a large human corpus will be wrong for one of them, so filtering has to be set per human-token budget.
- cost Following the advice to filter aggressively for human targets discards about 31% of an August 2026 crawl, leaving roughly 69% of collected tokens to train on.
- constraint Teams training models above about 180M parameters must refit the law or extrapolate past anything the paper tested.
Chinchilla-style scaling laws predict loss from model size and token count, and they price every token the same [7]. The new law keeps that shape but splits the data side into a benefit term and a harm term, so the marginal value of a token can change sign [8]. When the model is small or short on data, benefit dominates [8]. As it scales, harm takes over. The summary attributes the harm to bias and repetitive structure in AI text [15].
In budget terms, a fresh billion human tokens keeps lowering loss the way standard laws predict [16]. For a model that has already seen a large human corpus, a billion AI tokens can raise it [16]. Loss here means loss on human validation sets [5]. The summary's advice to filter aggressively is aimed at teams whose goal is human-centric tasks [13].
This is a different experiment from the collapse papers. According to the summary, earlier work such as Shumailov et al.'s "Curse of Recursion" fed a model only its own raw, unfiltered output [14]. Wild text comes from many models and is often edited by people [17]. It lands unlabeled in corpora like FineWeb, mixed into a far larger pool of human writing [12][17].
The experimental design deserves credit. The paper, "How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text", mapped the loss surface by pretraining 800 models across human-to-AI ratios [3][4]. The law was fitted on 50M-parameter models and used to predict models 3.6 times larger [9]. That puts its tested reach at about 180M parameters [1]. The post does not report the threshold ratio, the fitted coefficients, or an error figure behind its phrase "incredible precision" [9].
The corpus release is what lets another team refit the law at its own scale [10]. The summary says AI content was identified with the Pangram detector [11], and Pangram Labs is one of the two research groups [1]. A refit on that corpus inherits the detector's calls. A team that scores its own crawl with a different detector may get a different share to feed into the law.
The summary declares the "dead internet" essentially here [18]. By the study's own count, AI text was still under a third of August's tokens [1].
What to watch
- Whether the full paper publishes the threshold ratio and fitted coefficients, which would let a team compute an AI-text cutoff for its own human-token budget.
- A second detector's estimate of the AI share in the same August 2026 crawl, or an audit of the WildAI labels against it.
- A refit of the benefit-harm law at billion-parameter scale, showing whether the 3.6x extrapolation holds further out.