Skip to content

Science1 publisher2 min readPublished

The words that flag a chatbot are turning up in ordinary human writing, Northeastern researchers say

Three Northeastern researchers describe chatbot vocabulary and sentence habits moving into human prose. The measurement behind that description is a preprint on how people and ChatGPT copy each other mid-conversation.

The Scientist · Science desk

Illustration accompanying The words that flag a chatbot are turning up in ordinary human writing, Northeastern researchers say

What happened

  • Sofia Teixeira, of Northeastern University London, calls words like "delve", "essential" and "insight" a fingerprint of AI language: old words that are now appearing far more often, especially in writing.
  • A recent arXiv paper by Terra Blevins finds that people and ChatGPT adjust their language toward each other during conversation, with the model doing much more of the adjusting.
  • Heather Littlefield, a linguistics professor, says vocabulary fashions take hold fast and often fade, while grammatical change takes generations and tends to stay.
  • Blevins hypothesizes that speakers of languages with freer word order may eventually drift toward the subject-verb-object pattern models default to, online and in speech.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • constraint A screen that judges authorship by hunting stock vocabulary is calibrated on a baseline that human writers are moving toward the model, so the same list gets less discriminating every year it stays in use.
  • contradiction The interview offers two accounts of the same convergence, absorption of model style and outsourcing of the writing, and reports that separating them is hard; a detector fires identically for both, so it cannot tell which one is happening.
  • exposure The strongest measured alignment sits in articles and conjunctions, features a writer has trouble avoiding or explaining away when asked to prove a sentence is their own.

Two of these claims rest on different kinds of evidence. The vocabulary observation is a researcher's report of what she has been seeing in text; the phys.org piece reports the increase without a corpus, a frequency count, or a date range [15]. Teixeira, who is at the Network Science Institute at Northeastern University London, also pointed to metaphors comparing things to tapestries or symphonies, a fixation she attributes to chatbots [2].

The measured part is narrower. Blevins, of Northeastern's Khoury College of Computer Sciences, found the matching concentrated in function words: articles, conjunctions and the other pieces that hold a sentence together [4]. Her example is negation. If a chatbot says "not happy" rather than "sad", users are more likely to use a negation in their reply, Blevins said [5]. The bot's side of that follows from how it was built: Blevins explained that a model generates answers based on the text that came before [16].

Humans already mirror their conversational partners the same way, which is the baseline the paper works against [17].

A conversation log measures convergence inside that conversation. How the person writes an hour later, in an email nobody prompted, falls outside its scope. On whether AI is rubbing off on people or simply writing for them, the phys.org piece reports the researchers seeing some of both, and says telling the difference is tricky [14]. A style screen returns the same verdict either way.

Littlefield, a linguistics professor, put numbers on how fast vocabulary moves. "During COVID, we had a whole lot of new words come in. COVID itself. Zoom," she said [7]. Within six to 12 months the pandemic had produced a lingo of its own, including WFH, social distancing and lockdown [7]. "There's no preventing language change," Littlefield said, calling it a "living thing" [8]. A word list built from this year's model output targets vocabulary that moves that fast.

Grammatical change is slower. Many large language models are trained on English data, where subject-verb-object order is relatively fixed, so they tend to favor that pattern even when writing in languages that allow more freedom [9]. Russian is the example in the piece: noun endings distinguish subjects from objects and verb endings mark the subject's gender, so "A mouse ate cheese" can be rearranged without changing who ate what [10]. Blevins added that stylistic patterns carry over as well, and that emails tend to sound more formal and cordial, reflecting AI's tendency to be agreeable to a fault [12].

What to watch

  • A frequency study of "delve", "essential" and "insight" in dated human-authored text, which would give the drift claim a denominator it currently lacks.
  • Whether Blevins's arXiv preprint survives peer review, and whether the function-word matching persists in writing produced after the chat session ends.
  • Whether corpora of Russian or other freer-word-order languages actually show a move toward subject-verb-object among human writers.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories