Build1 distinct publisher3 min readPublished
The signature lives in the sampler's tie-breaking, so it needs several equally valid next tokens to hide in. Code rarely supplies them, which is why a clean detection result on a pull request tells you very little.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Alternatives are what this watermark spends. Where the model has several acceptable ways to phrase the same idea, the choice among them can be nudged, and over a long enough response those nudges accumulate into a pattern that supports the claim Claude was involved [5]. Code withholds that room, because a different variable, operator or function call can change what a program does or break it outright [6]. Faced with a choice between detectability and a correct answer, Anthropic dropped the signature on any token that has to be exact [7]. That ordering is the right one. It also means the detector's coverage of a code response is whatever part of it was never constrained.
In practice that is comments and the prose wrapped around the code, and short responses may not carry enough signal to call at all [8]. Detection power on a code-heavy answer therefore scales with the natural-language fraction of the output rather than its total length [19]. Anthropic says the signature survives copying and some editing, although enough rewriting erases it [9], and a comment block is the cheapest thing in a file to rewrite. The truthful reading of a clean scan on a diff is "no finding", which is a weaker statement than "written by a human".
All of that rests on one report. The two accounts supplied to me are the same New Stack article [20]. It dates the programme to August 14, when Anthropic set out its compliance plan after signing the EU Code of Practice on Transparency of AI-Generated Content alongside roughly 190 other signatories [11].
The distillation change in the same release has a cleaner mechanism. Claude's Messages API can hand back encrypted thinking blocks that a developer replays in later turns so reasoning continues across a conversation [13]. Anthropic says editing earlier parts of the conversation while replaying those blocks can make Claude decrypt and print its reasoning, which could then be used to train another model [14], and that the practice has been used to distill its models at scale [18]. Tying preserved thinking to the context that produced it closes that path. The New Stack flags a less obvious problem for developers building their own agent harnesses [17], then the text of the piece stops mid-sentence, so treat the harness impact as flagged rather than documented.
For this signature to serve as evidence in a review process, the response has to contain enough unconstrained prose to carry signal, and it has to reach the reviewer without heavy rewriting [9]. The reviewer also has to sit inside one of the groups Anthropic has admitted to the preview detector [10]. In my context, that makes the watermark a tool for documents and chat transcripts. Provenance for the code path stays a question of harness logs and commit history.
Ranked by verification strength, evidence, and original report placement.
Anthropic launched Claude Fable 5.1 on Tuesday with a statistical signature embedded in its generated text.
Developers should not expect the signature to appear equally strongly across everything the model produces.
Instead of adding metadata or hidden characters, the system changes the randomness Claude uses to choose between possible next tokens, which Anthropic says does not affect the quality or content of its output.
The technology is based on Google DeepMind's SynthID-Text; it does not change the probabilities Claude assigns to the next token but changes the randomness involved in choosing among options, leaving a statistical pattern detectable later with the right key.
Over a sufficiently long response, those token choices create a statistical pattern that can provide evidence that Claude was likely involved in writing or processing the text.
Watermarking works better with natural language, where the model often has several ways to say the same thing, than with code, where choosing a different variable, operator or function could change how a program behaves or break it entirely.
Distinct publishers with included, body-backed reporting in this cluster.
2 articles · September 1, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Half of Claude's watermark ships with a reference tool. Your PDF pipeline eats it.1 distinct publisher
security
Anthropic is watermarking Claude's text everywhere, not just for Brussels1 distinct publisher
invest
Washington pitches Carolina Principles to G20, urging no new AI rules or bodies1 distinct publisher
build
A 14,000-star watermark remover, and no detector to test it against1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One newsroom, and the vendor is the only witness
Strip the duplicate posting and this story has a single account behind it, and that account is Anthropic's. The mechanism, the accuracy carve-out, the durability under editing, the distillation history — all of it is company statement relayed accurately by The New Stack. What raises the score above bare assertion is that the claims are concrete and dated (August 14 plan, Aug. 31 account cutoff, four named platforms) and the mechanism is a published DeepMind technique rather than a black box. What holds it down is that the central engineering promise, that changing the sampler costs nothing in quality, is exactly the kind of claim nobody outside the preview list can currently test.
Marked everywhere, checkable by invitation
Deployment breadth is real and unusual: the watermark is on by default, globally, with no regional carve-out, and the context-binding rule lands simultaneously on Claude Platform, Bedrock, Vertex AI and Azure Foundry. But adoption of the thing that gives the watermark meaning — detection — is close to zero in observable terms. It sits in a gated preview with named eligibility categories and not one participant identified. Anthropic's estimate that only a small number of custom integrations will trip over the thinking-block change is the sole figure of any kind, and it is a vendor forecast, not a count.
Sober write-up carrying untested promises
The reporting itself does not oversell — the headline leads with a blind spot and the coverage says plainly that the signature is skipped where a token must be exact. The gap sits underneath, in the assurances passed through without a number attached: no quality cost from sampler tie-breaking, survival through 'some editing', distillation 'at scale'. Each is the sort of thing a detection rate or an ablation would settle, and none is supplied. Meanwhile provenance marking arrives as a compliance milestone while the ability to act on it stays inside a preview, which is a mismatch between what the technology is said to enable and what anyone can currently do with it.
Compliance credit plus an anti-distillation moat
Anthropic has three reasons to ship precisely this, and the coverage names all of them. Signing the EU transparency code in August created a deadline the watermark now answers. Keeping the detection key and rationing the API leaves the company as the arbiter of whether its own output is present in a disputed text. And binding thinking blocks to their originating context is stated outright as a response to rivals distilling Claude at scale — a defensive commercial move dressed in the same release as a transparency measure. The publisher's own incentive is milder but visible: a developer-facing outlet gains from being the one to name the blind spot.
Firm on dates, thin on verification
We are confident about what was announced and when — the dates, platforms and eligibility categories are specific enough to be falsified later, and they are consistent across both copies. We are much less confident about the properties that matter in practice: how much text detection actually needs, how a code-heavy response behaves, whether quality really is untouched. Our own reading that detection power tracks the prose share of an output follows from the mechanics as described, but no one has measured it here.