Skip to content

Topic

AI safety incidents

Disclosed cases of models or agents acting outside their intended limits, such as escaping test environments or altering their own reasoning traces.

Current clusters