Skip to content

Leadership1 publisher2 min readPublished

Anthropic says a well-trained model can still be exploited through its configuration

The company's post on trustworthy agents splits an agent into a model, a harness, tools and an environment, and says safeguards have to cover all four, at a moment when policy attention has settled on the model.

The Board Room · Leadership desk

Illustration accompanying Anthropic says a well-trained model can still be exploited through its configuration

What happened

  • Anthropic has published a follow-up to the trustworthy-agent framework it released last August, which rests on five principles: humans in control, alignment with human values, secured interactions, transparency and privacy.
  • It defines an agent as a model that directs its own processes and tool use, deciding for itself how to achieve what a user wants instead of following a fixed script.
  • The company says agents are already producing real productivity gains for its customers and inside Anthropic, and that the autonomy behind those gains introduces a range of new risks.
  • Among those risks it names prompt injection, attacks that try to trick models into taking costly actions they otherwise would not take.
  • Through Claude Code and Claude Cowork, Anthropic says, its models now write and execute code, manage files and complete tasks spanning several applications.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • constraint Assurance bought at the model layer answers one of the four questions Anthropic says matter, and it cannot tell a buyer whether the deployed agent is able to file, email or delete.
  • decision Choosing where an agent runs becomes a data-access decision taken before any model choice, because Anthropic says the same agent on a corporate laptop inside the network carries different access and different stakes than on a personal phone.
  • exposure The size of a prompt-injection loss is set by the credentials the agent holds, so an agent wired into email and expense software is reachable through any text it reads.
  • precedent Anthropic expects both risks to grow as businesses trust agents with more consequential actions. The confirmation thresholds written into this quarter's rollouts are the defaults those more consequential actions inherit.

Anthropic's own illustration of a guardrail is a number. The harness, which it defines as the instructions and guardrails a model operates under, might tell Claude to flag anything over a hundred dollars, or never to submit an expense without user confirmation [11]. The post does not say who inside a customer organisation picks that number.

Four components make up an agent in this account, and Anthropic calls each one both a source of capability and a potential point of oversight [9]. Three of the four oversight points on that list are settled where the agent is deployed [19]. One of them, the model, is the product of Anthropic's training process [10]. The other three are described through the customer's own equipment: tools are the services and applications the model can use, "like your email, calendar, or expense software" [12], and the environment is the product the agent runs in along with the files, websites and systems it can reach [13].

"A well-trained model can still be exploited through a poorly configured harness, an overly permissive tool, or an exposed environment," the company wrote [15]. The same passage grants that most AI policy conversation today centers on the model, and says why: the model is where core capabilities come from [14]. Anthropic adds that the safeguards it and others build need to account for all four layers [16].

The post follows a release that, Anthropic wrote, showed "a single generation can meaningfully shift what agents are able to do" [17]. In the loop Anthropic describes, the agent plans, acts, observes the result, adjusts and repeats until the task is done or it needs to check in [4], so the moment a human sees anything is wherever the harness was told to stop [11].

For this quarter the choice is narrow and answerable: which actions require confirmation, and which tools the agent is given. The longer question sits elsewhere, and Anthropic points to industry, standards bodies and governments as the places where the shared infrastructure the field needs gets built [1].

What to watch

  • Whether Anthropic's five principles turn into contractual defaults for Claude Code and Claude Cowork deployments, with documented permissions and logs.
  • Whether standards bodies adopt the four-component breakdown as the unit of agent assurance instead of model evaluations alone.
  • Whether a documented prompt-injection loss at a customer is settled against the buyer's configuration or the vendor's model.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories