build1 distinct publisher
Safety scores you can raise by saying no more often
A study with UK AI Security Institute researchers took eight safety benchmarks apart. The composite score is gameable, most questions are ballast, and over-cautious models leave traces.
Publishers:the-decoder.com
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+15
- Incentives45
- Confidence55