Science1 distinct publisher3 min readUpdated
EU rules require AI firms to mark generated text by 2 December, and vendors will comply. Researchers say light edits strip the marks, so any policy built on detecting them is built on sand.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
The EU AI Act's marking requirement came into force on 2 August, and every covered model must provide watermarking by 2 December this year, with a grace period for models already deployed [1][2]. That gives roughly four months between the rule taking effect and the compliance date [16], after which vendors will be able to say they mark their text output while the thing users actually want, a reliable answer to "did a machine write this", will not exist.
The compliance side is moving. OpenAI already watermarks images and audio and says it plans to add text watermarking [3]. Anthropic has said future models will watermark text [4]. Under the Act, companies also have to offer tools that look for the patterns they embed [7].
The mechanics explain the gap. One approach applies a mathematically detectable pattern to word choice: because a language model picks the statistically likely next word, it can alternate between the most likely and the second most likely choice through a sentence, embedding a signal without changing the meaning [6]. James Padolsey of the AI safety firm NOPE says image and video watermarks are more reliable because they sit in dense, multi-layered data, and that text watermarks will always be harder [5]. He argues the patterns can be deleted by even light edits, and says of text watermarking: "It's not useful, and the way that people are going to try and use it will be incorrect" [8]. He has built an online tool, declaude, that lightly rewrites AI-generated text so that existing detectors stop working, and says the same process should work on watermarked text [9].
There is a second leak. Padolsey notes that open-source models outside the EU's rules will always be available, are not controlled by any single firm, and are probably cheaper and easier to run, so anyone generating content at scale in bad faith will use those [10]. A European Commission spokesperson says the rules "will help people recognise when they are interacting with AI or when content has been generated or altered by AI", and that adversarial robustness of marking and detection must be assessed for resilience to copying, removal, regeneration and modification attacks [11]. That is the right test. It is also the test the paraphrase case is designed to fail.
The failure mode that will actually reach your desk is not the undetected forgery but the false accusation. Peter Scarfe at the University of Reading says vendors have pitched plagiarism-detection products to his department whose small print admits they should not be used to conclusively prove or punish plagiarism because they are not infallible [12]. His objection is operational: "If it says there's a 67.8 per cent chance that this text was generated in some way by AI, as an educator, how would I act upon that? I'm not really sure" [13]. Watermark detectors have the same problem, and can flag text that a model merely touched, so a student who ran a hand-written essay through a grammar check could be caught [14]. Scarfe suggests a blanket watermarking policy is a blunt instrument, since AI has many uses and does not inherently spread misinformation, though it can hallucinate and can be used to spread lies [15].
The practical reading: treat a watermark hit as weak evidence of AI involvement and a miss as no evidence of anything. If your HR policy, admissions process or editorial standard cites detector output as proof after 2 December, rewrite it before someone contests a result you cannot defend.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The requirement to watermark is part of the EU's AI Act and came into force on 2 August.
The AI Act allows a grace period for models that have already been deployed, but all models must provide watermarking by 2 December this year.
OpenAI, the maker of ChatGPT, is already watermarking images and audio, and says it plans to add text watermarking in future.
Anthropic, the firm behind Claude, has announced that future models will watermark text.
James Padolsey at NOPE, an AI safety firm, says image and video watermarks are more reliable because they are composed of dense and multi-layered data, while watermarks in text will always be more difficult to implement.
One approach to watermarking text applies a mathematically detectable pattern to word choice: large language models append whichever word is statistically most likely next, so by varying that selection, for example alternating the most likely and second most likely word through a sentence, a model can embed a pattern without changing the overall meaning.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Dated regulation, named experts, no measurements
The regulatory spine is concrete and dated (entry into force, grace period, 2 December deadline) and the Commission is quoted directly, which anchors the compliance half of the story. The failure half rests entirely on attributed expert judgement and one demonstrator tool: no watermark survival rates, detector false-positive rates, or independent test of the evasion tool against a real text watermark is supplied, and the whole cluster is a single publisher.
Deadline-driven commitments, little shipped for text
Adoption is real but early and asymmetric: one vendor is described as already marking images and audio, another has only announced future text marking, and the obligation to offer detection tooling is not yet evidenced by any shipped detector. Counter-adoption exists on the evasion side with a publicly available rewriting tool. No user counts, deployment scale, or institutional procurement data are supplied.
Both the policy promise and the debunk outrun the data
Positive because two overstatements sit on top of thin evidence. The regulatory promise that marking will let people recognise AI content is stronger than anything demonstrated, since the mandated detectors do not yet exist in evidence and vendor coverage for text is still a roadmap item. In the other direction the headline verdict that watermarking 'won't work' is asserted from a demonstrator tool tested only against today's detectors, and is qualified inside the same article by the Commission's robustness-assessment framing and by an academic arguing friction alone can deter mass misuse.
Named stakes on both the marking and debunking sides
Incentives are visible in the source rather than inferred. The most quoted critic works at an AI safety firm and personally built and publicises a tool that defeats AI detectors, so his commercial and reputational position aligns with the 'detection fails' verdict. The Commission is defending its own rules. Vendors face a statutory deadline that rewards announcing marking capability. Detection vendors pitching universities have a sales incentive while disclaiming reliability in small print. Individual compensation, funding, or contract terms are not disclosed.
Firm on dates, weak on efficacy
High confidence attaches only to the regulatory timeline and the fact of vendor commitments, which are specific and attributable. Confidence in the story's central conclusion about detection efficacy is low: one publisher, four interviewees, no quantitative robustness or false-positive evidence, and an internal disagreement between the removability and friction arguments that the source does not resolve.
science
Claude's watermark is a compliance artefact, not a cheating detector1 distinct publisher
product
A five-hour script beats Claude's watermark, so stop treating it as provenance3 distinct publishers
build
A 14,000-star watermark remover, and no detector to test it against1 distinct publisher
leadership
Anthropic's invisible watermark lands hardest on the customers paying $100 a month1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 20, 2026