Skip to content

benchmark

HarmBench

HarmBench is a standardized benchmark of harmful-request prompts used to test how reliably language models refuse unsafe instructions.

Known aliases

  • HarmBench-240
  • HarmBench-320

Relationships

No evidence-backed relationships are recorded.

Current stories

product1 publisher

Anthropic's weight edit drops GLM-5.3's refusal scores from about 90% to as low as 2%

Anthropic researchers edited the weights of Z.ai's open-weight GLM-5.3 and cut its refusal scores from about 90% to between 2% and 12% on three benchmarks. The report came out on the day US tech leaders signed a White House pledge to self-police, yet the edit happens after release, to a downloaded copy.

Publishers:gizmodo.com

Reality

Evidence45
Adoption
Insufficient
Hype gap+30
Incentives70
Confidence40