
New Scientist just published a piece on why firms are watermarking generated text, and why that is unlikely to work. I read it, and I want to be more direct than the headline. Statistical, token-pattern watermarking of the kind firms like Anthropic are recommending is the wrong approach. I give it a thumbs down.
The policy clock is real. The EU AI Act’s watermarking requirement came into force on 2 August 2026. Models already in the field get a grace period. Every model has to do this by 2 December. OpenAI already watermarks images and audio, and says text is coming. Anthropic says future models will watermark text. I understand why they are saying yes. I still think we are about to spend quality we actually need in order to satisfy a political mandate.
How the trick works
The method sounds tidy. A language model usually picks the most likely next token. To watermark, you bias that choice. One common sketch is to alternate the most likely token with the second most likely, so a detector can later see a pattern that a person would not notice. Meaning is supposed to stay the same. The detector is supposed to find the fingerprint.
That is the part I object to first. The whole point of a good model is that it picks the token that serves the user. If you force it to pick a worse token because the watermark needs a 1 instead of a 0, you have made the writing worse on purpose. Teachers, engineers, and anyone who relies on a careful sentence will feel that. We are throwing away useful generation quality so a detector can later nod.
The ownership problem
There is a second signal I do not like. A watermark can make the output look more owned by the model, more defensible as the model’s text. That is a bad ownership and control signal. I want people to treat generated text as a draft they are responsible for. A hidden pattern that says this belongs to the vendor works against that habit.
Anyone who wants to cheat can drive over it
James Padolsey at NOPE, quoted in the New Scientist piece, is blunt about the mechanics. Image and video watermarks sit in dense, multi-layered data. Text is thinner. Light edits strip the pattern. He built a tool called declaude that makes small changes so existing detectors fail, and he says the same idea will work on watermarked text. If you want to cheat with the tech, you will barely notice this.
Open-source models that sit outside EU rules will still be available. People who want unmarked text at scale will use those. They will often be cheaper and easier to run. The firms that follow the Act can check the box. The people the Act is most worried about will not.
Peter Scarfe at the University of Reading raises the other failure mode: false positives. His department keeps getting pitched plagiarism tools that, in the small print, say you should not punish anyone on the score. A 67.8 percent AI reading is not something an educator can act on. Watermark detectors will have the same shape. A student who only asked a model to check grammar could get flagged. That is a terrible classroom outcome, and it is a predictable one.
Friction is not a strategy
Erman Ayday at Case Western Reserve argues that even a trivial removal step creates friction. If you spend the time taking the mark off, maybe you could have written the thing yourself. I disagree that this saves the approach. The people generating slop at volume will automate the removal. The people writing carefully will pay the quality tax on every sentence. The mandate still gets its checkbox. The useful work gets worse.
A European Commission spokesperson said the rules help people recognize AI, and that robustness has to be assessed against removal and modification attacks. That last sentence is the whole problem. We already know light edits and open models defeat the scheme. Assessing robustness will not turn a doomed method into a good one.
What I would rather we do
Be honest with students and readers about when a model helped. Teach people to write and to check. Keep humans responsible for what they publish. If you need provenance, look at process and systems of record, not a hidden token dance that anyone can wash out.
Firms will still ship the watermark. They can still check the box. I will still give it a thumbs down.
Time to dig deeper
Start with New Scientist’s article. Then try the thing yourself. Take a paragraph you care about, run it through a model, then rewrite it by hand the way you would before sending it to a student or a client. Notice what you changed, and why. That exercise is more useful than a detector score. If you teach, write down how you would handle a 67.8 percent AI report before one lands on your desk. The better habit is still the same. Own the words you put your name on.