build1 distinct publisher
Item-level scoring splits WildJailbreak into a safety half and a reasoning half
Ai2 fit a multidimensional item response model to 100 models across 16 benchmarks. The per-question estimates say parts of its own safety suite are grading general reasoning rather than refusal behaviour.
Publishers:huggingface.co
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap−12
- Incentives55
- Confidence