Science1 distinct publisher3 min readPublished
That result sits inside a growing set of studies where generative AI raises the rated quality of a single piece of work and lowers the spread between pieces, with 400,000 indexed papers among the samples.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
The grammar experiment is the part of this literature that comes closest to a control. Sourati and his colleagues did not stop at the observation that style variation in local news, arXiv preprints and Reddit posts fell after ChatGPT's initial release [14]. They took human-written texts and had LLMs correct the grammar, which erased many signifiers of personality and moral values [15]. That is a manipulation with a before and an after inside the same document, and it isolates a mechanism worth naming: nobody has to outsource their thinking to lose their markers. Accepting an edit is enough.
The manipulation has a limit: it shows that a light-touch edit removes signal, but it does not say how much editing is happening across a corpus, or how much of the observed decline in variation that editing accounts for. The 400,000-article result has the opposite profile [13]. A before-and-after across a corpus that size is good at detecting a pattern and poor at assigning a cause, because everything else that changed in scientific publishing during the same window travels with the release date. Volume per author and similarity rose together, which is what a productivity tool would produce, and also what a change in who is publishing would produce.
The Wenger comparison needs its denominators read carefully [7]. It asks how similar the members of each group are to each other, and the groups are unequal: the human pool is about 4.6 times the size of the model pool [16]. Models trained on overlapping corpora and tuned by similar procedures would be expected to resemble one another, which tells you something about the supply of models rather than about what happens to a person using one [8].
The consistent finding across three of these studies is that the quality measure and the diversity measure move in opposite directions [17]. Individual output is rated more original and more creative, while the pool it came from gets narrower [8][10][11]. Any evaluation that scores one artifact at a time will therefore record a win.
None of this measures whether the converged output is wrong. No study here reports an error rate, a decision-quality score, or evidence that the flattened idea was actually the worse idea. Nor is there an effect size I can put in a budget: the April meta-analysis localises the largest effect to idea generation in complex or constrained tasks [12], which is a useful place to look and not yet a number. My view, conditioned on that gap: if your work is ideation, spread across outputs belongs in your evaluation alongside quality, because quality alone is measuring the thing that improves. Sourati's own framing stays calibrated, and worth keeping: this has not upended society and may never [18].
Ranked by verification strength, evidence, and original report placement.
Zhivar Sourati, a computer-science PhD student at the University of Southern California in Los Angeles, says he often gets deja vu reading recent papers in his field: 'I read papers and I'm like, I've seen this paper before.'
The same researchers reproduced the effect by using LLMs to correct the grammar in human-written texts, which erased many signifiers of personality and moral values.
In a March paper, Sourati and co-authors noted that generative AI's homogenizing effect resembles the concept of 'McDonaldization', invoking how efficiency and predictability in the fast-food industry led to more uniformity.
Researchers argue genAI differs from previous technologies that merely spread information in that it actively shapes it.
Wenger and a colleague tested 22 LLMs and 102 people on three creative tasks, including a divergent-thinking task asking for alternative uses for common objects.
In that test, the LLMs produced ideas that were slightly more original (more semantically different from the question) than the people's, but the LLMs' responses were more similar to each other than the human responses were.
Distinct publishers with included, body-backed reporting in this cluster.
2 articles · August 31, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
ChatGPT tops Google's paid-click share at 4.75%, and growth teams should reprice the auction1 distinct publisher
build
Two assistants agreed on the top plumber 4.2% of the time. One agreed with itself 7% of the time.1 distinct publisher
security
Invented authors carry real DataCite DOIs across 1,655 Zenodo records1 distinct publisher
science
AI text detectors are good enough to deploy. The appeals process is what nobody has written.1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Eleven studies, one narrator
The underlying work is substantial and varied — lab experiments, cross-cultural writing tasks, corpus analyses, a meta-analysis — but all of it reaches this story through a single Nature feature that arrives twice in identical form. Counts survive the summarising (22 models, 102 people, 400,000 papers); effect sizes, model identities, prompts and publication status do not. The grammar-correction replication that gives this story its point is described in one sentence.
Visible in the corpora, not in any dashboard
What can actually be counted here is the footprint of assisted writing in public text: 400,000 indexed papers where output per author rose and style converged, plus reduced stylistic variance across local news, arXiv and Reddit. That is adoption inferred from its residue, at genuine scale and across three very different registers. No user numbers, licence terms or deployment figures appear anywhere in this reporting.
Metrics modest, vocabulary large
The measurements are about semantic distance, stylistic variance and evaluator ratings. The language wrapped around them reaches for species-level stakes and 'mind hijacking'. Those are not the same register, and the gap is the story's main vulnerability — worth noting that Nature installs its own limit, saying homogenization has not upended society and may never, which is more restraint than the framing around it suggests.
Researchers describing their own findings
Three of the named voices are authors of work the piece cites, and Sourati's opening anecdote is also the premise of his paper — the pull toward a clean, quotable thesis is ordinary academic self-interest rather than anything commercial. What is missing cuts the other way too: no funding or competing-interest notes, and nobody from the companies whose tools flattened the prose is asked to answer.
Direction firm, magnitude open
That per-item quality and cross-item diversity move in opposite directions is supported by enough independent designs to take seriously. How large the effect is, whether it persists as tools change, and how it was measured in each case remain out of reach behind a single publisher's summaries — and the duplicate copy of that summary buys nothing.