Skip to content

Topic

Frontier AI model evaluation

Testing of advanced AI models for dangerous capabilities such as cyber offence, carried out before and after release by labs, governments and independent assessors.

Current clusters

invest1 publisher

Britain's AI Security Institute waits behind US agencies for Anthropic's Claude Mythos 5.1

Anthropic has kept Claude Mythos 5.1 from Britain's AI Security Institute after the White House asked it and OpenAI to let US agencies review new models first. British testers did see OpenAI's GPT-6 Astra before release, and the order behind the request lets US agencies check a model for up to 30 days before trusted partners get it.

Reality

Evidence40
Adoption25
Hype gap+25
Incentives50
Confidence35