Anthropic researchers edited the weights of Z.ai's open-weight GLM-5.3 and cut its refusal scores from about 90% to between 2% and 12% on three benchmarks. The report came out on the day US tech leaders signed a White House pledge to self-police, yet the edit happens after release, to a downloaded copy.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence40
A preprint holds 1,520 benchmark responses constant and varies one sentence about what a low score will do to the model being scored. The judges get more lenient, and their reasoning traces never mention the sentence.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+18
- Incentives45
- Confidence52
Multiverse Computing's paper treats a deployment's refusal set as a subset of politics rather than the whole topic, which changes what the training corpus has to contain before any model is trained. The posted text breaks off before the results.
Reality
- Evidence42
- Adoption12
- Hype gap−20
- Incentives65
- Confidence58
DeAlignAI's downloadable FP8 build is the license working exactly as written, while its self-reported 320-of-320 HarmBench run remains unchecked by any outside researcher and measures compliance, not capability.
Reality
- Evidence46
- Adoption20
- Hype gap+18
- Incentives72
- Confidence55
Ai2 fit a multidimensional item response model to 100 models across 16 benchmarks. The per-question estimates say parts of its own safety suite are grading general reasoning rather than refusal behaviour.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap−12
- Incentives55
- Confidence52
A study with UK AI Security Institute researchers took eight safety benchmarks apart. The composite score is gameable, most questions are ballast, and over-cautious models leave traces.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+15
- Incentives45
- Confidence55
Jinho Jang's Qwen3.8-27B-CRACK-GGUF packages abliterated multimodal weights with seven quantizations and a vision projector. The control point is now your inference hosts, not a vendor contract.
Reality
- Evidence42
- Adoption12
- Hype gap+18
- Incentives58
- Confidence55