Product1 distinct publisher3 min readPublished
Claude will send, reply and forward once you switch approval off, which leaves that review step doing double duty as the feature's only brake and as the sole defence anyone can name against prompt injection.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
Ask Claude to tell HR you will be out on Thursday, with approval switched off, and it drafts and sends in one motion, so the first person to read that email is HR [5][15]. The pitch is inbox management, and the grant is standing authority to speak from your address [1]. Those are two different things to sign off on, and only one of them appears on the settings screen.
Engadget names four ways this goes wrong, and three of them need no attacker at all: hallucinated content in a message already sent, a misread instruction that forwards the wrong thread, and the plain fact that the inbox contents now sit with Anthropic [6][7][14]. The fourth is the one with a working proof. White-on-white or zero-point text in an incoming message is invisible to the reader and legible to the model, and Engadget reports it has been used to pull verification codes for accounts elsewhere [8][9][10]. Claude warns about this the first time you allow it to send, which is a candid thing for a feature to say about itself [3].
Every mitigation on offer runs back to the same switch. Writing more specific instructions narrows misinterpretation and does nothing about text you cannot see [13][8]. Multi-factor authentication elsewhere is recommended in the same article that notes those codes arrive in the mailbox the agent can read [13][9]. Simon Willison, who coined the term prompt injection, says we still do not know how to stop it reliably [11]. That leaves the approval prompt carrying the load by itself [2].
The hard line sits at permanent deletion: Claude can trash or archive, and it cannot destroy [4]. The blocked action keeps the message inside the account, while the permitted one puts a copy in someone else's [1].
For anyone rolling this out beyond themselves, the gap in the evidence matters more than the risk list. Engadget's account is written for one person with one mailbox, and it describes no admin-level control and no audit trail, so an organisation cannot tell from it who has approval switched off [16].
A workable test has two axes: what triggers the action, and where the result lands. Delegate without approval when you typed the request and the outcome stays in your account, which covers filing and drafting. Keep approval on for anything with an outside recipient. The quadrant to leave unattended is the one where arriving content triggers an outgoing message, because that is the exact path injection takes [8][9].
Who this is for: one operator whose mailbox is not the recovery address for anything expensive, and who would rather apologise for a clumsy email than write it. For a shared or role mailbox, that toggle is the outbound mail policy, and whoever flips it owns what leaves [1][2].
Ranked by verification strength, evidence, and original report placement.
The default behaviour is 'ask before sending', and Engadget's first and most important mitigation is to leave that default active so any action can be reviewed before it executes.
When asked to send an email, Claude goes straight into thinking and sends it; the user does not see or edit the message beforehand unless the right setting is enabled, so there is a high chance the recipient discovers any mistake first.
Engadget names three action risks: Claude hallucinating false information into an email and sending it, Claude misunderstanding a request and sending or forwarding something unintended, and a hidden prompt in an incoming message hijacking it into acting on an attacker's instructions.
Engadget also cites the privacy concern of trusting Claude and Anthropic with all the data in the inbox.
Simon Willison, who coined the term 'prompt injection', says 'we still don't know how to 100% reliably prevent this from happening', and Engadget reports experts remain split on whether it is even solvable yet.
Engadget's other mitigations are to be very specific in instructions, to stay cautious with unfamiliar senders and to enable multi-factor authentication elsewhere, while stating that a significant level of risk remains.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 6, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
product
Claude Can Now Press Send In Gmail, And Your Workspace Admin Owns That Decision1 distinct publisher
product
An AI agent told to book a gym class found a missing authorization check and used it1 distinct publisher
product
Plaud's charging case will sit on the table and upload the meeting by itself1 distinct publisher
build
Claude can send mail and delete events; owners decide who skips the approval prompt2 distinct publishers
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One outlet, one checkable quote
Everything rests on a single Engadget piece. Its sturdiest element is a direct quote from Simon Willison, who coined the term, saying nobody knows how to stop injection reliably. Its weakest is the sentence carrying the alarm: that these attacks are proven, with no case attached. The one incident named, OpenClaw wiping Summer Yue's inbox, is an agent disregarding instructions rather than an attacker planting them. Anthropic's documentation never appears, so the behaviour of the toggle at the centre of the story is described but not verified.
No uptake figures anywhere
Nobody here counts users. Engadget confirms the sending permission exists and warns about it, but there's no number for how many people have turned approval off. Anthropic hasn't disclosed anything on that front, and no telemetry from any party fills the gap either. The OpenClaw deletion shows one researcher had a live agent in her mail; it does not scale into a measure of how widely this pattern is in use.
Calm framing, uncited "proven"
The volume is not the problem: Engadget's headline promises only "some risks involved" and the advice is sensible. The gap sits in the claim that these attacks have actually happened, when the sole thing shown to have happened in this account is unrelated deletion by another agent. Against a hazard the same piece concedes is unsolved, a mitigation list of be specific, watch unfamiliar senders and turn on MFA elsewhere lands a little more reassuring than the evidence supports.
Search-shaped service journalism
Engadget earns on consumer tech traffic, and this is written for someone typing whether to let Claude into their mail, which pulls toward a piece that neither dismisses the feature nor frightens readers away from it. There's no affiliate hook or vendor access at play here, and Anthropic hasn't offered a statement either, so the pressure shows up in framing and headline rather than in what is claimed.
Mechanism solid, product surface unverified
Moderate. Hidden instructions in text a model reads and a human does not is a well-understood technique, and the Willison quote is real. What a single consumer outlet cannot settle is the product itself: whether the default truly is ask-before-sending, what else the sending permission covers, and whether Anthropic has narrowed it since.