Science1 publisher3 min readPublished
Text watermarks land on 2 December. The detection they imply does not.
EU rules require AI firms to mark generated text by 2 December, and vendors will comply. Researchers say light edits strip the marks, so any policy built on detecting them is built on sand.
The Scientist · Science desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The requirement to watermark is part of the EU's AI Act and came into force on 2 August.
- The AI Act allows a grace period for models that have already been deployed, but all models must provide watermarking by 2 December this year.
- OpenAI, the maker of ChatGPT, is already watermarking images and audio, and says it plans to add text watermarking in future.
- Anthropic, the firm behind Claude, has announced that future models will watermark text.
- James Padolsey at NOPE, an AI safety firm, says image and video watermarks are more reliable because they are composed of dense and multi-layered data, while watermarks in text will always be more difficult to implement.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
The EU AI Act's marking requirement came into force on 2 August, and every covered model must provide watermarking by 2 December this year, with a grace period for models already deployed [1][2]. That gives roughly four months between the rule taking effect and the compliance date [16], after which vendors will be able to say they mark their text output while the thing users actually want, a reliable answer to "did a machine write this", will not exist.
The compliance side is moving. OpenAI already watermarks images and audio and says it plans to add text watermarking [3]. Anthropic has said future models will watermark text [4]. Under the Act, companies also have to offer tools that look for the patterns they embed [7].
The mechanics explain the gap. One approach applies a mathematically detectable pattern to word choice: because a language model picks the statistically likely next word, it can alternate between the most likely and the second most likely choice through a sentence, embedding a signal without changing the meaning [6]. James Padolsey of the AI safety firm NOPE says image and video watermarks are more reliable because they sit in dense, multi-layered data, and that text watermarks will always be harder [5]. He argues the patterns can be deleted by even light edits, and says of text watermarking: "It's not useful, and the way that people are going to try and use it will be incorrect" [8]. He has built an online tool, declaude, that lightly rewrites AI-generated text so that existing detectors stop working, and says the same process should work on watermarked text [9].
There is a second leak. Padolsey notes that open-source models outside the EU's rules will always be available, are not controlled by any single firm, and are probably cheaper and easier to run, so anyone generating content at scale in bad faith will use those [10]. A European Commission spokesperson says the rules "will help people recognise when they are interacting with AI or when content has been generated or altered by AI", and that adversarial robustness of marking and detection must be assessed for resilience to copying, removal, regeneration and modification attacks [11]. That is the right test. It is also the test the paraphrase case is designed to fail.
The failure mode that will actually reach your desk is not the undetected forgery but the false accusation. Peter Scarfe at the University of Reading says vendors have pitched plagiarism-detection products to his department whose small print admits they should not be used to conclusively prove or punish plagiarism because they are not infallible [12]. His objection is operational: "If it says there's a 67.8 per cent chance that this text was generated in some way by AI, as an educator, how would I act upon that? I'm not really sure" [13]. Watermark detectors have the same problem, and can flag text that a model merely touched, so a student who ran a hand-written essay through a grammar check could be caught [14]. Scarfe suggests a blanket watermarking policy is a blunt instrument, since AI has many uses and does not inherently spread misinformation, though it can hallucinate and can be used to spread lies [15].
The practical reading: treat a watermark hit as weak evidence of AI involvement and a miss as no evidence of anything. If your HR policy, admissions process or editorial standard cites detector output as proof after 2 December, rewrite it before someone contests a result you cannot defend.