InvestNot yet confirmed elsewhere1 publisher3 min readPublished
Physicists' formula for when chatbots turn bad got 15 of 16 small-model tests right
George Washington University physicists published a formula predicting when chatbots turn bad, after it called 15 of 16 preprint test cases. That evidence comes from six small open models and 300-token chats, a narrow base for the offline safety monitor the authors propose.
The Investor · Invest desk

What happened
- The study appeared in the journal Patterns and builds on a preprint the pair first posted in February, before formal peer review.
- Johnson and Huo trace the failures to the attention head, which a growing conversation's context pulls toward one cluster of possible answers until it tips.
- The formula estimates n, the count of good tokens before the first bad one, and n is zero when a conversation already leans toward the bad side.
- To push the tipping point out of reach, the authors describe injecting content into a conversation so it lands beyond the length of the response.
- According to the authors, alignment training can shift or suppress tipping for specific prompts but cannot remove its underlying cause.
Why it matters
- capability Offline deployments, cut off from the cloud services most safety tools depend on, get a check that can in principle run on the device beside the model.
- cost If training only moves tipping prompt by prompt, an offline team's safety budget has to pay for a runtime check on the device on top of the training run.
- exposure Users of offline companion chatbots carry the risk if a monitor calibrated on sub-billion-parameter models misjudges the more capable models Decrypt expects on devices as hardware improves.
Fifteen of 16 is 93.75%, rounded to 94% in Decrypt's account, and it is one miss [14]. The 16 were described as clear-cut cases. Each test was a binary call on whether a model would tip at once or after a run of good tokens [3]. The preprint's predictions could also be off by one output [9].
The six open-weight models in the preprint came from OpenAI, EleutherAI and Meta and ran from 124 million to 410 million parameters [8], so the largest was about 3.3 times the smallest [15]. The published paper reportedly extends the tests to seven models of up to 12 billion parameters [17]. That is about 29 times the preprint's ceiling [19]. Decrypt still calls it small by current standards [17], and its account does not give a hit rate for the larger set. The preprint's window was 300 tokens, a few short paragraphs [9]. A chat that tips after its 300th token was outside what those tests could measure [16].
The authors have in mind a phone or laptop running a model with no internet connection, companion chatbots included [7]. On that hardware their monitor runs in parallel with the model and flags when the predicted tipping point, n*, drops below a safety threshold [10]. The authors call the monitor low-cost [10]. An app builder would test that claim first, because the compute comes out of the same device that is running the chatbot.
This can play out a few ways. If the 12-billion-parameter tests land near 15 of 16, an offline team gets a warning it can compute locally. That matters because, as the authors argue, existing safety tools often depend on a cloud connection these models lack [6][10]. If the hit rate falls as models grow, the formula fits small models and loses value just as Decrypt expects on-device models to get more capable with better hardware [18]. A third possibility is that the formula holds but running the monitor in parallel costs more on a phone than the authors expect.
We think the work is strong enough for an on-device team to build a prototype monitor around. It is too thin for anyone to quote 94% as a deployment accuracy. The counter-case is that the authors tie tipping to how the attention head weighs accumulated context [4], which is a claim about how the model works and not about its size. The pair have also moved from one deliberately simplified attention head in their April 2025 paper [13] to six downloadable models with a single miss [8][14]. If the paper reports a hit rate for the 12-billion-parameter set well below 15 of 16, the case for prototyping is wrong. A rate at or above it would mean we were too cautious.
What to watch
- The hit rate the Patterns paper reports for its seven models of up to 12 billion parameters, set against the preprint's 15 of 16.
- Tests on conversation windows longer than 300 tokens, where a late tip would show up.
- Any on-device chatbot app that ships the parallel n* monitor and publishes what it costs to run.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+30
- Incentives
- Insufficient
- Confidence45
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
George Washington University physicists Neil Johnson and Frank Yingjie Huo published a formula that estimates how many good tokens an AI model produces before its first bad one.
- [2]
The study appeared in the journal Patterns and builds on a preprint, posted publicly before formal peer review, first released in February.
- [3]
In the preprint, the formula picked the right case, immediate or delayed tipping, in 15 of 16 clear-cut tests, or 94%.
- [4]
Johnson and Huo trace the problem to the attention head; as a chat grows, accumulated context pulls attention toward one cluster of possible answers or another until it tips.
- [5]
The formula estimates the tipping point n as the number of good tokens before the first bad one; if the conversation already leans bad the model tips right away with n of zero, and if it leans good it delivers a run of fine answers then flips.
- [6]
The authors argue that existing safety tools often depend on a cloud connection that offline models lack.
- [7]
The target is on-device AI that runs entirely on a phone or laptop with no internet connection, including companion chatbots.
- [8]
The preprint tests ran on six open-weight models from OpenAI, EleutherAI and Meta, all between 124 million and 410 million parameters.
- [9]
The preprint's tests used small models and a 300-token window, or a few short paragraphs of text, and its predictions could be off by one output.
- [10]
The authors propose a low-cost monitor that runs in parallel with the model and flags when n* falls below a safety threshold.
- [11]
Alignment training can shift or suppress tipping for specific prompts but cannot remove the underlying mechanism, the authors say.
- [12]
The authors describe ways to push the tipping point out of reach, such as injecting content into the conversation so n* lands beyond the length of the response.
- [13]
An earlier paper from the same pair, covered by Decrypt in April 2025, modeled a single, deliberately simplified attention head.
- [14]
15 correct calls out of 16 is 93.75%, leaving one miss.
- [15]
The largest preprint model was about 3.3 times the size of the smallest.
- [16]
A conversation that tips after its 300th token fell outside what the preprint's tests measured.
- [17]
The published paper reportedly widens the test to seven models of up to 12 billion parameters, which Decrypt describes as still small by current standards.
- [18]
Decrypt says on-device AI is a trend that may grow as hardware becomes more powerful and smaller AI models become more capable.
- [19]
The published paper's largest model, at 12 billion parameters, is about 29 times the preprint's 410-million-parameter ceiling.
Sources
1 independent publisher whose own reporting we read for this story.
- decrypt.coHere’s a Way to Predict When AI Chatbots Will Turn Bad
1 article · October 10, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Transformer attentionFollow
- On-Device AIFollow
- AI safetyFollow