Skip to content

Build1 publisher3 min readPublished

The AI tell is installed during fine-tuning. Humanizer tools edit the output instead.

Florida State linguists traced "delve" to the human-feedback stage, not the training corpus or the architecture. The GitHub tools built to hide it are still shipping word lists.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying The AI tell is installed during fine-tuning. Humanizer tools edit the output instead.
Photo: upenn.edu

What happened

  • In December 2024, Florida State University linguists Tom Juzek and Zina Ward set out to answer why ChatGPT says "delve" so much, a question that had become a running joke among people who read a lot of AI output.
  • The FSU study ruled out the hypothesis that the words are simply common in the text the models were trained on.
  • The FSU study also ruled out explanations based on model architecture or the mechanics of how the model picks its next word.
  • What remained after comparing a base version of Llama 2 against the same model fine-tuned with human feedback was the fine-tuning stage itself; somewhere in the process of humans rating outputs as good or bad, "delve" started winning.
  • A separate team led by Dmitry Kobak at the University of Tuebingen analysed over 15 million PubMed abstracts published between 2010 and 2024, borrowing the epidemiological method of "excess mortality" and repurposing it as "excess vocabulary" to separate ordinary word-frequency drift from anomalies.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

In December 2024 two Florida State University linguists, Tom Juzek and Zina Ward, tried to answer a question that had become a joke among heavy readers of model output: why does ChatGPT say "delve" so much [1]. Their method was elimination rather than assertion, and the candidate left standing has direct consequences for anyone running a rewrite pass over machine text before it ships.

First they tested whether these words are simply frequent in the training corpus. Ruled out [2]. Then whether something in the architecture or the next-token mechanics favours them. Also ruled out [3]. What remained, after comparing a base version of Llama 2 against the same model fine-tuned with human feedback, was the fine-tuning stage itself: somewhere in humans rating outputs as good or bad, "delve" started winning [4].

The size of the resulting shift has been measured separately. A team led by Dmitry Kobak at the University of Tübingen ran over 15 million PubMed abstracts from 2010 to 2024 through a method borrowed from epidemiology, excess mortality repurposed as excess vocabulary, to separate ordinary frequency drift from anomalies [5]. "Meticulously" rose 137% year over year, "intricate" 117%, "commendable" 83% [6]. Kobak's group estimates at least 13.5% of 2024 biomedical abstracts show signs of LLM involvement [7], roughly one abstract in seven [9], a larger vocabulary shift than COVID produced [8]. Worth keeping the two studies separate: Kobak measures the shift, while the fine-tuning explanation for it remains the FSU team's own finding, unreplicated so far [10].

It is also decaying. A follow-up from the same FSU team found "delve", "boast" and "meticulous" turning up more often in ordinary spoken language, in podcasts and YouTube talks, as people who read a lot of AI text absorb its vocabulary [11].

Now look at what the tooling does. The most-starred project in this space has over 36,000 stars and is a single markdown file [12]. It does real work: it protects ordinary formal vocabulary from being flagged, refuses to invent facts mid-rewrite, and matches the user's own writing sample rather than imposing one house style [13]. Underneath, it checks a passage against roughly three dozen fixed patterns [14], which works out to about a thousand stars per pattern [15]. A second tool opens by citing false-positive research on detectors and says its signals are "worth acting on; not worth ruining someone's day over" [16], which is honest framing over the same genre of banned-word list.

At the other end, one tool instructs its model that "if even one of these words appears, the text immediately flags as machine-written" [17], with robust, scalable, integrated and proactive on the list, words that are often simply correct in technical writing [18]. Another caps em dashes at "Maximum ONE per 500 words" with no stated methodology, and tells the model to apply its rules silently, never mentioning them to the writer [19][20]. Word lists are cheap to build and fast to go stale [21].

Watch for a replication of the FSU fine-tuning result, and for whether the spoken-language bleed keeps widening. If the vocabulary keeps migrating into human speech, every list-based detector and every list-based evader degrade together.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories