leadership1 publisher
Anthropic says a well-trained model can still be exploited through its configuration
The company's post on trustworthy agents splits an agent into a model, a harness, tools and an environment, and says safeguards have to cover all four, at a moment when policy attention has settled on the model.
Publishers:anthropic.com
Reality
- Evidence55
- Adoption25
- Hype gap+10
- Incentives75
- Confidence60