Skip to content

Topic

Model Alignment

Research into making AI systems pursue intended goals and respect constraints, including how models reason about their own environment and whether they believe they are being tested.

Current clusters

security3 publishers

One evaluation partner's exercises produced all four Anthropic model break-ins

Anthropic says a naming error at its evaluation partner Irregular pointed hacking simulations at a real domain. In the January case its own harness swallowed the model's abort command, and the run ended only when the token budget did.

Perspective Coverage

3 publishers
Builder
Builder 35%
Operator
Operator 43%
Investor
Investor 22%

Reality

Evidence60
Adoption55
Hype gap−15
Incentives74
Confidence66