Skip to content

Topic

Model Collapse

Degradation of models trained recursively on their own output, and published findings on when it does and does not occur.

Current stories

build1 publisher

Wild AI text, about 31% of web tokens, gets a harm term in a new scaling law

Pangram Labs and UMass Amherst researchers put AI-generated text at about 31% of August 2026 web tokens and fit a scaling law for when it starts to hurt. For teams that pretrain on web crawls, AI text becomes a measured line item in the data budget, with a curve for its marginal value.

Publishers:dev.to

Reality

Evidence35
Adoption
Insufficient
Hype gap+40
Incentives45
Confidence40