Probes on Qwen3.5-27B reading other models' text came within 0.004 AUROC, on average, of probing authors up to 397B directly, a LessWrong post reports. Every tested pair was open-weight, so auditors who apply the method to closed models get the reader's view of the text and cannot measure that gap.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence30
The Oxford team told two agents driven by one model to count cards, and the agents worked out the rest, putting bet signals inside chatter about the dealer. Catching them needed activations from inside both models at once.
Reality
- Evidence45
- Adoption15
- Hype gap+30
- Incentives60
- Confidence40
A new 27-family benchmark scores how well a tool reads a model's intermediate variables during a forward pass. Its authors say it does not fully rule out tools that infer the answer from the prompt.
Reality
- Evidence40
- Adoption10
- Hype gap+10
- Incentives55
- Confidence45
A replication of Wurgaft et al.'s manifold steering on Anthropic's pre-trained CLT features reproduces the cyclical weekday transition in Gemma-2-2B. The author attributes the weaker push to the MLP sublayers holding only part of the signal.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap−5
- Incentives25
- Confidence45
Almost all frontier protein models ship with open weights, out of reach of API refusals and unlearning. A LessWrong essay puts the enforcement point at DNA synthesis screening, where AI-designed novel sequences already defeat the sequence matching that providers use.
Reality
- Evidence33
- Adoption15
- Hype gap+18
- Incentives45
- Confidence44
Anthropic screened Claude for one property, verbalizability, and found two more in the same representations. Neel Nanda reproduced the structure in open weights, and the outside commentators Anthropic invited disagree about what it is.
Reality
- Evidence40
- Adoption22
- Hype gap+12
- Incentives62
- Confidence42
Wired reports that Jacob Coxon's September 8 exit post from Anthropic was followed by a senior colleague putting internal odds of human extinction at 10 percent, and by Dario Amodei's weekend essay arguing for slower releases.
Reality
- Evidence34
- Adoption28
- Hype gap+26
- Incentives66
- Confidence38
Base Labs has signed Hugging Face and Goodfire to publish methods for training, evaluating and monitoring open models, with the check running where the activations are. The partners have not published an evaluation suite yet.
Reality
- Evidence42
- Adoption10
- Hype gap+38
- Incentives72
- Confidence48
The offset in a J-lens readout mostly tracks how often a token appears, and scaling it by variance, after plain subtraction failed, lifted hidden-word elicitation to 0.805 from 0.665 on Gemma-2-9B-it. The paired test over 20 words gives p of about 0.19.
Reality
- Evidence46
- Adoption12
- Hype gap+10
- Incentives40
- Confidence56
A LessWrong post shows the Oja rule falling out of tied sparse-autoencoder descent once two terms are dropped, then reports a language-model test where an initialization change moved the backprop baseline more than the Hebbian gap it was meant to explain.
Reality
- Evidence46
- Adoption
- Insufficient
- Hype gap+12
- Incentives34
- Confidence51
In a single-author LessWrong experiment, accuracy on held-out countries went from 20.9% to 97.7% with the lesion active in every forward pass, and the recovery survived refitting the lens to the adapted model.
Reality
- Evidence34
- Adoption12
- Hype gap+22
- Incentives18
- Confidence42
A one-dimensional probe finds structural impossibility in the hidden state of instruction-tuned models from 1.7B to 70B parameters. The trained refusal pathway that guardrail work tunes reads an axis about 85 degrees away from it.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap−12
- Incentives40
- Confidence52
Perturbation probing needs the weights, which puts the finding on teams running open-source checkpoints rather than on API tenants. The same 50-neuron toolkit that breaks a refusal template also repairs other behavior.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+30
- Incentives80
- Confidence45
A STAT opinion piece argues the measurable shift in medical AI is not benchmark wins over physicians but the removal of the safety text that used to route people to care.
Reality
- Evidence48
- Adoption71
- Hype gap+22
- Incentives74
- Confidence44
Knowledge editing is sold as a cheap substitute for retraining. The authors argue it treats models as filing cabinets when knowledge is a dependency graph.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+22
- Incentives58
- Confidence45