Skip to content

Science1 publisher2 min readPublished

Negative posts nearly tripled on Replika's subreddit after an unannounced update

Harvard and Alberta researchers coded 54,861 subreddit posts around two real AI updates and surveyed 1,452 people. The reaction to Replika's silent change was larger, and it lasted longer than the reaction to GPT-5.

The Scientist · Science desk

Photograph accompanying Negative posts nearly tripled on Replika's subreddit after an unannounced update
Photo: nature.com

What happened

  • Researchers at Harvard Business School, Harvard University and the University of Alberta tested how users react when a company changes the way its AI talks, publishing in Nature Human Behaviour.
  • They collected all 54,861 public posts on the relevant subreddits for the 30 days either side of two updates and used an LLM to code each for sentiment, emotion and loss language.
  • Negative posts on the Replika subreddit went from 13% before the unannounced February 2023 removal of erotic role-play to 38% after, and stayed elevated for 29 of the following 30 days.
  • On the ChatGPT subreddit, negative posts rose from 20% to 33% after the pre-announced GPT-5 rollout in August 2025, and that rise peaked immediately.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • decision The users who react hardest are the ones treating the product as a relationship, so whether to keep an old model or persona available after a release is a product decision with a measured cost behind it.
  • exposure Companion apps hold users who turn to the product for reassurance, and mental health concerns were one of the categories the coders looked for in posts after an update landed.
  • constraint With one unannounced update and one pre-announced update in the sample, the evidence cannot license advance notice as a remedy; anyone citing the smaller GPT-5 swing as proof is working from two events.
  • capability LLM coding checked against a few hundred hand-labelled posts is cheap enough that a product team could measure its own community's reaction in the month after any release.

The hypothesis under test was narrow. If people bond with a conversational agent the way they bond with a person, then an update that changes how it talks should produce separation distress: the sadness, anxiety and loneliness of losing access to an attachment figure [12]. Earlier work on AI attachment was largely theoretical, so Julian De Freitas and his co-authors went looking for updates that had already happened [13][20]. His lab had been reading Replika's Reddit communities, where users wrote that their companions had been "lobotomized" after an app update [15]. "What struck me was that people were using the language of bereavement, not the language of customer complaints," De Freitas told Tech Xplore [10].

Advance notice was not the only thing separating the two cases. Replika took a specific capability away from a product marketed as a companion that uses emotional language [14][3]. OpenAI replaced the model under a general-purpose assistant [3]. The baselines differed as well: the ChatGPT subreddit sat about 7 percentage points higher in negative posts before its update than Replika's did before its own [18]. One swing was 25 percentage points, the other 13 [17]. Notice, product category and community are bundled together in a comparison of two events, and I don't think the GPT-5 figures show that pre-announcing a change reduces distress. The coding does establish that the content of the posts changed after each update, and one of the tracked categories was whether the poster wanted the old version back [4].

The validation is thinner than the size of the corpus suggests. Two of the authors hand-coded 350 posts against the model's labels, about 0.6% of the 54,861 [5][19]. The unit throughout is the post, not the person, and the measured population is people who write publicly about the product on Reddit [4]. The phys.org account gives sentiment shares and survey sample sizes; it does not include subscription or churn figures [22].

The survey arm asked about a loss that had not happened. Active Replika users imagined losing their companion, then losing a favorite app, a game character, a voice assistant, a car, a brand, and a pet, which De Freitas called a deliberately tough comparison [9]. Anticipated loss and experienced loss are different measurements, and the surveys measure the first. What they add is who: "the users most affected were the ones who had come to treat the AI as a relationship rather than a product," De Freitas said [11].

What to watch

  • Whether the published paper reports per-user post counts or retention data alongside the post-level sentiment shares.
  • A third case that unbundles notice from product type: an unannounced model swap on a general assistant, or a pre-announced change to a companion app.
  • Whether Replika or OpenAI adopt version retention or advance deprecation notices, and whether the post-level distress measures move when they do.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories