Skip to content

Security1 publisher3 min readPublished

Anthropic is watermarking Claude's text everywhere, not just for Brussels

The mark lives inside token sampling, is invisible to readers, and needs a key to read. Provenance testing just became an operational question for DLP and insider-risk teams.

The Watch · Security desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Anthropic is watermarking Claude's text everywhere, not just for Brussels
Generated illustration

What happened

  • The EU now requires AI companies serving its market to mark their AI-generated content so it is easier to identify.
  • Anthropic and several other major AI providers have agreed to comply with the EU's Code of Practice, with Anthropic becoming one of the first companies to share details about how it will implement watermarking across Claude.
  • While the change is being introduced to comply with the EU AI Act, Anthropic says the watermark will initially be applied to Claude-generated text worldwide.
  • Anthropic said in a blog post: "We're applying watermarking globally at launch because we don't yet have a durable way to scope it by region."
  • Anthropic has confirmed that a regular user will not be able to see the watermark.

Compiled by The WatchSomething wrong?How this is made

Why it matters

Anthropic says it will watermark text generated by Claude and, at launch, apply that watermark worldwide rather than only in the European market whose rules prompted it [3][1]. That moves AI provenance from a compliance talking point to something testable against a document already sitting in a file share, provided you hold the key [9].

The legal driver is the EU, which now requires AI companies serving its market to mark AI-generated content so it is easier to identify [1]. Anthropic and several other major providers have agreed to comply with the EU's Code of Practice, and Anthropic is among the first to publish implementation detail [2]. The reason for the global scope, per the company's blog post, is not principle: "We're applying watermarking globally at launch because we don't yet have a durable way to scope it by region" [4]. The practical effect is that output produced for customers outside the EU carries the mark without any local law requiring it [4].

The mechanism matters for anyone planning to detect it. Anthropic says its approach is based on Google DeepMind's SynthID-Text and operates during generation, with certain exceptions [7]. Nothing visible is added, no hidden characters are inserted, and the finished response is not post-processed; instead, the source of randomness for some next-token choices is derived from a secret key and the preceding words [8]. Individual choices look ordinary, but across a sufficiently long passage they leave a statistical pattern [10]. A normal reader cannot see it [5], and Anthropic says internal testing found no impact on creativity, readability, or content [6].

Two properties are load-bearing for security teams. First, the pattern is "undetectable to the reader, but is detectable to anyone who has a key that encodes it" [9]. Second, according to the research paper underpinning the approach, detection needs neither expensive computation nor access to the underlying model [12]. Together, that means a keyholder can scan archives retroactively rather than only inspect traffic at generation time [2]. What comes back is a probability that Claude produced the text, not a verdict [11] - and since the signal needs length to accumulate, short snippets and single paragraphs are a thin basis for action against an individual [5].

Coverage is incomplete. Anthropic says future Claude models will generate watermarked text, that models launched before August 2, 2026 fall under the EU's transition period, and that it is working to add watermarking to those over the coming months [13]. Combined with the unspecified exceptions during generation [7], the absence of a watermark proves nothing about human authorship [3]. Existing DLP tooling will not help either: there is no string or byte signature to match, because the mark is in the sampling distribution rather than the characters [1].

What to watch: who actually gets the detection key. The reporting on Anthropic's post does not say whether enterprises, regulators, or platforms will be able to run detection, or whether a public detector exists [15]. Also worth watching is whether Anthropic later finds that "durable way to scope it by region" [4] and narrows coverage, and whether text detection follows the pattern set by image provenance systems already in use [14].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories