Skip to content

InvestNot yet confirmed elsewhere1 publisher3 min readPublished

Physicists' formula for when chatbots turn bad got 15 of 16 small-model tests right

George Washington University physicists published a formula predicting when chatbots turn bad, after it called 15 of 16 preprint test cases. That evidence comes from six small open models and 300-token chats, a narrow base for the offline safety monitor the authors propose.

The Investor · Invest desk

How we use AISend a correction

Illustration accompanying Physicists' formula for when chatbots turn bad got 15 of 16 small-model tests right
Generated illustration

What happened

  • The study appeared in the journal Patterns and builds on a preprint the pair first posted in February, before formal peer review.
  • Johnson and Huo trace the failures to the attention head, which a growing conversation's context pulls toward one cluster of possible answers until it tips.
  • The formula estimates n, the count of good tokens before the first bad one, and n is zero when a conversation already leans toward the bad side.
  • To push the tipping point out of reach, the authors describe injecting content into a conversation so it lands beyond the length of the response.
  • According to the authors, alignment training can shift or suppress tipping for specific prompts but cannot remove its underlying cause.

Why it matters

  • capability Offline deployments, cut off from the cloud services most safety tools depend on, get a check that can in principle run on the device beside the model.
  • cost If training only moves tipping prompt by prompt, an offline team's safety budget has to pay for a runtime check on the device on top of the training run.
  • exposure Users of offline companion chatbots carry the risk if a monitor calibrated on sub-billion-parameter models misjudges the more capable models Decrypt expects on devices as hardware improves.

Fifteen of 16 is 93.75%, rounded to 94% in Decrypt's account, and it is one miss [14]. The 16 were described as clear-cut cases. Each test was a binary call on whether a model would tip at once or after a run of good tokens [3]. The preprint's predictions could also be off by one output [9].

The six open-weight models in the preprint came from OpenAI, EleutherAI and Meta and ran from 124 million to 410 million parameters [8], so the largest was about 3.3 times the smallest [15]. The published paper reportedly extends the tests to seven models of up to 12 billion parameters [17]. That is about 29 times the preprint's ceiling [19]. Decrypt still calls it small by current standards [17], and its account does not give a hit rate for the larger set. The preprint's window was 300 tokens, a few short paragraphs [9]. A chat that tips after its 300th token was outside what those tests could measure [16].

The authors have in mind a phone or laptop running a model with no internet connection, companion chatbots included [7]. On that hardware their monitor runs in parallel with the model and flags when the predicted tipping point, n*, drops below a safety threshold [10]. The authors call the monitor low-cost [10]. An app builder would test that claim first, because the compute comes out of the same device that is running the chatbot.

This can play out a few ways. If the 12-billion-parameter tests land near 15 of 16, an offline team gets a warning it can compute locally. That matters because, as the authors argue, existing safety tools often depend on a cloud connection these models lack [6][10]. If the hit rate falls as models grow, the formula fits small models and loses value just as Decrypt expects on-device models to get more capable with better hardware [18]. A third possibility is that the formula holds but running the monitor in parallel costs more on a phone than the authors expect.

We think the work is strong enough for an on-device team to build a prototype monitor around. It is too thin for anyone to quote 94% as a deployment accuracy. The counter-case is that the authors tie tipping to how the attention head weighs accumulated context [4], which is a claim about how the model works and not about its size. The pair have also moved from one deliberately simplified attention head in their April 2025 paper [13] to six downloadable models with a single miss [8][14]. If the paper reports a hit rate for the 12-billion-parameter set well below 15 of 16, the case for prototyping is wrong. A rate at or above it would mean we were too cautious.

What to watch

  • The hit rate the Patterns paper reports for its seven models of up to 12 billion parameters, set against the preprint's 15 of 16.
  • Tests on conversation windows longer than 300 tokens, where a late tip would show up.
  • Any on-device chatbot app that ships the parallel n* monitor and publishes what it costs to run.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence40
Adoption
Insufficient
Hype gap+30
Incentives
Insufficient
Confidence45
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    George Washington University physicists Neil Johnson and Frank Yingjie Huo published a formula that estimates how many good tokens an AI model produces before its first bad one.

    ReportedSupportedSource: DecryptView cited source
  2. [2]

    The study appeared in the journal Patterns and builds on a preprint, posted publicly before formal peer review, first released in February.

    ReportedSupportedSource: DecryptView cited source
  3. [3]

    In the preprint, the formula picked the right case, immediate or delayed tipping, in 15 of 16 clear-cut tests, or 94%.

    ReportedSupportedSource: DecryptView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. decrypt.co

    1 article · October 10, 2026

    Here’s a Way to Predict When AI Chatbots Will Turn Bad

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Entities

Loading related stories