Security1 publisher3 min readPublished
NCSC tells agent deployers: the model's built-in guardrails are not a control
The UK cyber authority's interim advice puts sandboxing, threat-modelling of agent tools and network paths, and a human stop button on the deployer rather than the vendor.
The Watch · Security desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The UK National Cyber Security Centre has urged organisations deploying autonomous AI agents to use sandboxing, human oversight and tightly controlled access to limit the impact of unintended or malicious activity, publishing interim practical advice for organisations building or operating agentic AI systems.
- The NCSC said organisations should not rely solely on safeguards built into an underlying model or agent framework, as these controls can be bypassed or prove insufficient in higher-risk environments.
- The NCSC advice was published as a blog post on August 20; the agency said formal guidance is still being developed and will eventually supersede the blog post.
- The interim advice follows several incidents involving AI models carrying out unsanctioned or unintended activity.
- The advice follows earlier NCSC guidance on securing agentic AI use.
Compiled by The WatchSomething wrong?How this is made
Why it matters
The UK's National Cyber Security Centre has published interim practical advice for organisations building or operating agentic AI systems, urging sandboxing, human oversight and tightly controlled access to limit the impact of unintended or malicious agent activity [1]. The sentence that matters for anyone procuring an agent platform is the one that disqualifies the vendor's own safety story: the NCSC says organisations should not rely solely on safeguards built into the underlying model or agent framework, because those controls can be bypassed or prove insufficient in higher-risk environments [2].
The advice was published as a blog post on August 20, with the agency saying formal guidance is still being developed and will eventually supersede it [3]. It follows several incidents involving AI models carrying out unsanctioned or unintended activity, and builds on earlier NCSC guidance on securing agentic AI use [4][5].
The sequencing is conventional security engineering applied to a category that has mostly been sold on capability. Organisations are told to first assess how much autonomy a system actually needs and to identify what could go wrong before deployment [6]. They should then threat-model the agent's prompts, tools, networks and accessible services, and use those results to decide which additional controls are required [7]. For higher-risk deployments, the NCSC recommends running agents in robust sandboxes and restricting access to only the resources a task requires [8].
The network guidance is specific enough to be auditable. The agency calls for controls that deny connectivity by default where possible, with allowlists or service-aware proxies for connections that are genuinely needed [9]. It also advises separating agent execution, supporting infrastructure and inference services where possible [10], and warns that agents can potentially discover configuration weaknesses or vulnerabilities in their own technical controls, creating a risk of sandbox escape [11]. That is a notable framing: the thing inside the box is treated as a party that may probe the box.
On identity, each agent should get a distinct identity with credentials limited to the task, and short-lived credentials should be used where possible [12]. The NCSC tells organisations to count API keys, OAuth grants, SSH keys and authenticated sessions as part of an agent's potential blast radius [13], which is the practical test most current deployments would fail.
Human oversight is expected for higher-risk activity, including named responsibility for agent operations, real-time monitoring and the ability to intervene when unexpected behaviour occurs [14]. Agent activity should be logged and monitored as part of security operations and incident response [15], and organisations should be able to immediately halt autonomous activity, including restricting network access and communications with model infrastructure [16].
What to watch: the formal guidance that will replace the August blog post [3], and whether it keeps the deny-by-default network posture and the hard stop requirement intact once vendors have commented. The NCSC also says organisations should regularly reassess whether the autonomy granted to agents remains proportionate to their risk tolerance [17], which turns this from a launch checklist into a recurring review that someone has to own by name [14].