Build1 distinct publisher3 min readUpdated
Florida State linguists traced "delve" to the human-feedback stage, not the training corpus or the architecture. The GitHub tools built to hide it are still shipping word lists.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
In December 2024 two Florida State University linguists, Tom Juzek and Zina Ward, tried to answer a question that had become a joke among heavy readers of model output: why does ChatGPT say "delve" so much [1]. Their method was elimination rather than assertion, and the candidate left standing has direct consequences for anyone running a rewrite pass over machine text before it ships.
First they tested whether these words are simply frequent in the training corpus. Ruled out [2]. Then whether something in the architecture or the next-token mechanics favours them. Also ruled out [3]. What remained, after comparing a base version of Llama 2 against the same model fine-tuned with human feedback, was the fine-tuning stage itself: somewhere in humans rating outputs as good or bad, "delve" started winning [4].
The size of the resulting shift has been measured separately. A team led by Dmitry Kobak at the University of Tübingen ran over 15 million PubMed abstracts from 2010 to 2024 through a method borrowed from epidemiology, excess mortality repurposed as excess vocabulary, to separate ordinary frequency drift from anomalies [5]. "Meticulously" rose 137% year over year, "intricate" 117%, "commendable" 83% [6]. Kobak's group estimates at least 13.5% of 2024 biomedical abstracts show signs of LLM involvement [7], roughly one abstract in seven [9], a larger vocabulary shift than COVID produced [8]. Worth keeping the two studies separate: Kobak measures the shift, while the fine-tuning explanation for it remains the FSU team's own finding, unreplicated so far [10].
It is also decaying. A follow-up from the same FSU team found "delve", "boast" and "meticulous" turning up more often in ordinary spoken language, in podcasts and YouTube talks, as people who read a lot of AI text absorb its vocabulary [11].
Now look at what the tooling does. The most-starred project in this space has over 36,000 stars and is a single markdown file [12]. It does real work: it protects ordinary formal vocabulary from being flagged, refuses to invent facts mid-rewrite, and matches the user's own writing sample rather than imposing one house style [13]. Underneath, it checks a passage against roughly three dozen fixed patterns [14], which works out to about a thousand stars per pattern [15]. A second tool opens by citing false-positive research on detectors and says its signals are "worth acting on; not worth ruining someone's day over" [16], which is honest framing over the same genre of banned-word list.
At the other end, one tool instructs its model that "if even one of these words appears, the text immediately flags as machine-written" [17], with robust, scalable, integrated and proactive on the list, words that are often simply correct in technical writing [18]. Another caps em dashes at "Maximum ONE per 500 words" with no stated methodology, and tells the model to apply its rules silently, never mentioning them to the writer [19][20]. Word lists are cheap to build and fast to go stale [21].
Watch for a replication of the FSU fine-tuning result, and for whether the spoken-language bleed keeps widening. If the vocabulary keeps migrating into human speech, every list-based detector and every list-based evader degrade together.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
According to the dev.to analysis, word lists share the same fragility no matter how carefully they are built: cheap to build and fast to go stale.
In December 2024, Florida State University linguists Tom Juzek and Zina Ward set out to answer why ChatGPT says "delve" so much, a question that had become a running joke among people who read a lot of AI output.
The FSU study ruled out the hypothesis that the words are simply common in the text the models were trained on.
The FSU study also ruled out explanations based on model architecture or the mechanics of how the model picks its next word.
What remained after comparing a base version of Llama 2 against the same model fine-tuned with human feedback was the fine-tuning stage itself; somewhere in the process of humans rating outputs as good or bad, "delve" started winning.
A separate team led by Dmitry Kobak at the University of Tuebingen analysed over 15 million PubMed abstracts published between 2010 and 2024, borrowing the epidemiological method of "excess mortality" and repurposing it as "excess vocabulary" to separate ordinary word-frequency drift from anomalies.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Named studies with specific figures, but one publisher and unnamed tools
The research side is comparatively well specified: named researchers and institutions, a base-versus-fine-tuned Llama 2 comparison, a 15-million-abstract PubMed corpus, and concrete percentages. The source also self-limits by flagging the causal fine-tuning result as unreplicated. The tooling side is weaker: every critiqued repository is anonymous, so the 36,000-star figure, the 'three dozen patterns' count and the quoted rules cannot be checked from the supplied material, and there is only one publisher in the cluster.
Real usage on both sides: heavy tool popularity and measurable LLM text in published science
Adoption is evidenced in two directions. Demand for evasion tooling is visible in a 36,000-plus-star repository, and the underlying volume of machine-assisted writing is quantified at 13.5% or more of 2024 biomedical abstracts, a shift the source calls larger than COVID's. Diffusion of AI vocabulary into podcasts and talks suggests further spread. Scores are held short of high because stars are a popularity proxy rather than deployment, and no install, traffic or organisational-use data is supplied.
Popular tools and a viral causal story both outrun their evidence
Two overstatements sit in this cluster. First, the widely repeated 'fine-tuning causes the AI tell' explanation rests on a single unreplicated team's finding, while the well-measured PubMed work speaks only to magnitude. Second, the tooling's implied capability — undetectable, humanized prose — is delivered by three dozen fixed patterns, absolute single-word flags and an undocumented em-dash quota, and the marker itself is already leaking into human speech. The gap is positive but moderate, because the article's own framing is deflationary rather than promotional.
No incentive disclosures in supplied material
The supplied source contains no funding, sponsorship, employment, vendor-relationship or commercial-interest disclosures for the author, the critiqued repositories, or the cited research teams. GitHub stars are reported as popularity, not as revenue or a business model, and no detector vendor is named. Nothing here supports a scored incentive reading without inference.
Directionally credible, single-source and partly unverifiable
Confidence is moderate-low. The direction of the story is credible and internally consistent, and the author is candid about the limits of the causal finding. But the cluster has one publisher, no primary documents, anonymised tooling, and its keystone mechanism claim is unreplicated — enough to report cautiously, not enough to treat as established.
build
AI-written code fails the same four ways, and every gate you own reports green1 distinct publisher
science
Text watermarks land on 2 December. The detection they imply does not.1 distinct publisher
build
A 5x publishing increase cost one site 1,000 indexed pages and every impression1 distinct publisher
build
48 startups, 4 known by name, 28 recommended by category1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 18, 2026