Published Build3 min read
Anthropic's Watermark Answers The Wrong Question, And Its Own Docs Say So
The Claude watermark survives light editing and dissolves under heavy editing, and its absence proves nothing. Policies that treat it as a verdict are built on a binary that does not exist.
Written for builders.See today for builders

What happened
- Anthropic signed the EU AI Act's Code of Practice on Transparency of AI-Generated Content and started marking text produced by Claude with an invisible statistical watermark.
- Anthropic's official documentation states the watermark does not prove a text was entirely AI-generated: a human can write a text, have Claude correct or translate it, and the result can still carry the mark.
- The reverse also holds: no watermark detected does not prove a human wrote everything.
- The watermark is a statistical signal baked into the model that openly acknowledges its limits: it can disappear after enough editing or translation, and its presence says nothing about who came up with the idea in the first place.
- A Medium article called the Anthropic announcement a "nuclear bomb" for the AI-generation community and claimed "you are effectively carrying a digital scarlet letter", without citing the actual documentation.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Anthropic has signed the EU AI Act's Code of Practice on Transparency of AI-Generated Content and begun marking text produced by Claude with an invisible statistical watermark [1]. The mark works as advertised; the problem is that it does not answer the question editorial and compliance teams will actually be handed, which is not "was a model involved" but "who is answerable for this text".
Read the documentation and the limits are stated openly. Per Anthropic's own materials, as reported by dev.to writer sylwia-lask, the watermark does not prove a text was entirely AI-generated: you can draft something yourself, have Claude correct or translate it, and the output still carries the mark [2]. The converse also holds. No watermark detected does not prove a human wrote everything [3]. The signal can also fade after enough editing or translation, and its presence says nothing about who originated the idea [4].
So the watermark is a two-way false negative and a two-way false positive with respect to authorship. That has not stopped the reading of it as a verdict. One Medium post called the announcement a "nuclear bomb" for the AI-generation community and told readers they are "effectively carrying a digital scarlet letter", without citing the documentation at all [5].
Pascal Cescato, writing on dev.to, points out that this collapses three unrelated mechanisms into one causal chain: Anthropic's watermark, third-party consumer detectors such as ZeroGPT, and platform decisions like a displayed badge or a curation algorithm that penalises flagged content [6]. The detectors predate Anthropic, have nothing to do with it, and according to Cescato their reliability has never been seriously demonstrated at scale [7]. The badges and ranking penalties are editorial choices made platform by platform, not a mechanical consequence of the mark [8].
The detector layer is where the damage lands, because that is what most platforms and clients will actually run. Cescato took an article he published in April 2021, before any consumer-facing LLM existed, and ran it through ZeroGPT twice: roughly a year ago it scored 97% AI probability, and a recent run on the identical text with the same tool returned 8.6% [10]. That is an 88.4 point swing on a text that did not change by a single mark of punctuation [11]. In the second run, the flagged passages were the most neutral and pedagogical ones: a definition of a web server, an explanation of CentOS Stream, a step-by-step automated update procedure. The passages carrying his voice, including self-deprecation, verbal tics and a Neapolitan moka pot bought at a flea market, were never flagged [12]. One test proves nothing general, as he says, but it suggests the tool responds to stylistic neutrality and structural regularity rather than origin [13].
The distinction worth building policy on is the one Cescato draws: assisted text stays under human editorial control end to end, generated text comes out of a prompt without that upstream control, and industrial-scale production is an automated pipeline with no editorial oversight at all. What separates them is who kept their hand on the decisions, not how much the model intervened [9]. Cescato notes he still has no answer for the middle case: a text thought through, written and reviewed by a human whose final English phrasing came out of a model [14].
Watch what platform disclosure rules actually cite when they land: the watermark, a third-party detector score, or a declared editorial responsibility. The first two are measurements of the wrong variable. Watch also whether anyone publishes a threshold, because a threshold on a signal that swung 88 points on unchanged text is a policy written against noise [11].
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Anthropic signed the EU AI Act's Code of Practice on Transparency of AI-Generated Content and started marking text produced by Claude with an invisible statistical watermark.
ReportedView cited source - [2]
Anthropic's official documentation states the watermark does not prove a text was entirely AI-generated: a human can write a text, have Claude correct or translate it, and the result can still carry the mark.
ReportedSource: Anthropic documentation as read and reported by dev.to writer sylwia-lask, summarised by Pascal CescatoView cited source - [3]
The reverse also holds: no watermark detected does not prove a human wrote everything.
- [4]
The watermark is a statistical signal baked into the model that openly acknowledges its limits: it can disappear after enough editing or translation, and its presence says nothing about who came up with the idea in the first place.
ReportedView cited source - [5]
A Medium article called the Anthropic announcement a "nuclear bomb" for the AI-generation community and claimed "you are effectively carrying a digital scarlet letter", without citing the actual documentation.
- [6]
Cescato identifies three mechanisms that are conflated in the panic reading: Anthropic's watermark, third-party consumer detectors such as ZeroGPT, and platform decisions such as a badge displayed on Medium and a curation algorithm that penalises flagged content.
Sources & coverage · 3 publishers
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- the-decoder.comMatthias BastianAug 14Anthropic announces watermark detection API that will let third parties detect Claude's AI texts
- dev.toPascal CESCATOAug 15The "AI" Badge Doesn't Measure What You Think It Does
Cited in this coverage: Anthropic documentation as read and reported by dev.to writer sylwia-lask, summarised by Pascal Cescato
Cited in this coverage: Anthropic documentation as reported by sylwia-lask on dev.to

