Build1 publisher3 min readPublished
Half of Claude's watermark ships with a reference tool. Your PDF pipeline eats it.
Anthropic's text mark is a keyed scheme no third party can test. The C2PA credentials on images since 11 August are checkable by anyone, until a PDF generator strips them.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Half one of Anthropic's watermarking announcement is text: a keyed statistical mark spread across token choices, which the author says an individual cannot test.
- Anthropic began attaching C2PA content credentials to generated images on 11 August.
- C2PA is an open standard with an open reference tool, so this half is checkable by anyone today with no key and no vendor cooperation.
- On 14 August, in the comment thread of Sylwia Lask's post, the author noted that Anthropic updated its documentation and Claude is using SynthID; Anthropic's page describes it as a version of the SynthID-Text approach published by Google DeepMind.
- The scheme has a published specification: Nature 634, 818-823.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Anthropic's watermarking announcement has two halves, and only one of them is a thing your team can act on this week. The text half is a keyed statistical mark that a third party cannot verify at all [1][9]; the file half, C2PA content credentials attached to generated images since 11 August, is an open standard with an open reference tool, checkable today with no key and no vendor cooperation [2][3].
Start with the half that is settled. The author of the dev.to post reports that on 14 August, in the comment thread of Sylwia Lask's original piece, its author noted Anthropic had updated its documentation: Claude uses SynthID, described by Anthropic as a version of the SynthID-Text approach published by Google DeepMind [4]. That moves the mechanism out of speculation and into a specification, published as Nature 634, 818-823 [5]. The scheme hashes the last few tokens with a secret key and uses that seed to bias which of several equally good next words is emitted, so in unmarked text those tokens arrive at chance rate and counting them yields a z-score and a real p-value [6]. A calibrated false-positive rate is the actual advantage over a classifier [7].
It also means the published attacks apply. Re-tokenisation, which can be as simple as inserting a character between every word and deleting it again, destroys the n-gram contexts a hash-based scheme seeds from, and OpenAI names that attack in its own writing on why it never shipped text watermarking [11]. Paraphrase removes the mark outright below roughly 800 tokens [12]. Jovanovic, Staab and Vechev at ICML 2024 showed an attacker can both spoof and scrub state-of-the-art schemes for under $50, with over 80% average success [13]; spoofing, forging a mark onto human text, is the direction that hurts a person and that nobody plans for [14].
Two housekeeping myths, per the same post. The mark is not zero-width Unicode: not U+200B, not U+FEFF, not the tag block [8]. So stripping invisible characters from a paragraph does nothing to a statistical watermark, because the signal is in which words were chosen [15]. And finding invisible characters proves nothing either, since they arrive from copy-paste, CMSs and PDF extraction [16]. The author notes their own text extraction strips those characters first for prompt-injection reasons unrelated to watermarking [10].
The gate is symmetry. Without the secret key you cannot reconstruct the seeds, and without the seeds there is nothing to count, so for a third party holding a suspicious paragraph the test cannot be run [9]. The fallback is worse than its reputation: Liang et al. found seven detectors flagged 61.22% of human-written TOEFL essays as AI, and 78.3% of granted patent claims were flagged [17][18].
Which leaves the file half as the only half with an operator story, and that is where the author's measurement lands. On 14 August they put a signed image through every PDF generator on their laptop to see what a downstream reader could still see, having built a PDF signal engine [19][20]. The angle here is unambiguous: most PDF generators strip that provenance on the way out. Note the load-bearing gap. The source excerpt describes the experiment and its motivation but the per-generator results are not in the material available to me, so treat the specific pass/fail table as unpublished here rather than as a result I can cite.
What to watch: whether Anthropic ever offers a verification endpoint for the text mark, since without one the EU AI Act's machine-detectability requirement is satisfied for the vendor and nobody else [21]. And whether PDF toolchains start preserving C2PA manifests, because if they do not, the only checkable half of this announcement dies at the export step.