The mark on the wrong word

The mark on the wrong word

Posted on: 21 August 2026

The European regulation asks model providers to mark what their systems produce. The instrument built to satisfy that request can only mark the places where the writing did not matter, and the gap between those two sentences is the whole story.

Before going further I should say how this piece was made, because it is evidence rather than disclosure. I brought the subject to Claude, the model checked the facts and corrected a premise I had got wrong, the technical inversion the argument rests on surfaced while it was reading the original documentation, and I then chose which of the angles offered was worth keeping and rewrote the text because in the shape it reached me I would not have put my name to it. Anyone who has read me before can hear where the machine stops. That is not a boast, it is a difference of rhythm.

Article 50 of the EU AI Act became applicable on 2 August 2026, requiring providers of generative models to mark their outputs in machine-readable form. Around a hundred and ninety signatories put their names to the Code of Practice on Transparency of AI-Generated Content in July. Penalties for non-compliance run to fifteen million euros or three per cent of worldwide annual turnover, whichever is greater. Anthropic answered on 14 August with a page explaining how the mark works in Claude's text, and the explanation is considerably more interesting than the announcement.

Nothing is added to the text. No hidden characters, no metadata riding alongside, nothing that disappears when you paste into a plain editor. A language model generates one word at a time, choosing among candidates, and a great many of those choices are indifferent: after "the weather was cold and" the next word can be grey or overcast without the sentence meaning anything different, and in cases like that the choice is settled by a random number. The watermark changes only where that randomness comes from. Instead of an arbitrary generator it uses a cryptographic key combined with the preceding words, so the choices remain unpredictable to a reader while anyone holding the key can test whether the sequence is consistent with the decisions the model would have made using it. The method is a version of SynthID-Text, published by Google DeepMind in Nature in 2024 and descended from a proposal Scott Aaronson made in 2022. Anthropic explains it through a game of Monopoly in which the players, rather than rolling dice, take their moves from the decimal digits of pi starting at some arbitrary point. The play is as random as it ever was for anyone sitting at the board, and afterwards someone who knows pi can work out that the dice were never used. It is a better analogy than these things usually get and worth keeping for other purposes.

What follows from the mechanism is the part almost nobody is discussing. The mark can only live where an indifferent choice exists. Where the right word is the only word there is nothing to mark: a proper name, a date, a figure, the exact title of a work, a line of code that breaks if you swap a term. Anthropic gives the example of Isaac Newton's most famous work being called Principia, where the next word has to be Mathematica and no other. In factual passages the signal thins until it vanishes. In discursive passages, where the possible rephrasings are many and all of them acceptable, it thickens. The same holds for proofreading. If the text is mine and the changes are few, the mark has too little to attach itself to and simply fails to register.

So the instrument measures lexical choice and stays blind to where the thinking came from. Someone who has a model build the argument, verify the sources and assemble the figures, then writes it up in their own hand, comes out clean. Someone who works out an original position alone and asks for help making it readable comes out marked. Anthropic states the limit with a clarity it deserves credit for: the mark cannot distinguish "Claude wrote this" from "Claude heavily edited this". The public argument about plagiarism assumes a moral order in which taking the substance is serious and having the surface tidied is venial, and the first official instrument of measurement inverts that order exactly. It records the surface. On the substance it is blind by construction rather than by imperfection.

The temptation at this point is obvious and I had it first, so it should be named. The temptation is to use the asymmetry to build oneself an exemption: I am not one of the people who copy, I use it to make the work better, my case is different. It is a weak move. Anyone making it has already conceded that a tribunal exists and is merely asking to be tried under the right heading, which hands the tribunal everything it needs. The plea of exemption confirms the legitimacy of the criterion it was meant to contest.

The defensible position costs more. In analytical writing the authorship does not sit in the words. It sits in what you decide to assert, in what you are prepared to be wrong about in public, and in what you refuse to publish when the facts will not hold it. The watermark measures none of that, and not because it was built badly. It cannot, for the same technical reason it cannot mark a date.

The real difficulty is not suspicion but its distribution. Anthropic says that a complete rewrite in which every word is replaced removes the mark, while light editing probably does not. The constraint is therefore free to evade for anyone who understands the mechanism and fully borne by anyone who does not. A rule with that property does not change behaviour, it sorts populations: those who have read the technical documentation carry on exactly as before, those who have not pay the entire bill. Something close to this happened to sampling after Grand Upright Music v Warner in 1991, when Judge Kevin Duffy opened his decision by quoting the seventh commandment. Nobody banned sampling. It became an activity requiring a legal department, and a compositional method that had been available to anyone with a sampler was left to those who could afford to clear every fragment. British records of the period that had been built out of dozens of borrowed seconds stopped being made, not because the technique was outlawed but because the paperwork cost more than the record.

There is a calendar detail worth noticing as well. The mark applies to models released from 2 August onwards, while for earlier ones the Act allows a transition period and the retrofit will arrive over the coming months. For some unspecified stretch of time, then, marked and unmarked texts coexist, produced by the same provider by the same method. Anyone treating the absence of a mark as proof of anything will be talking nonsense, and will do it anyway.

The detection interface is not yet available. When it arrives, the first public accusation built on a watermark will say more than any analysis: if it lands on someone who copied a whole text, the system does what it says. If it lands on someone who had a paragraph of their own rewritten, we know what has been built and we know it will go on working that way.


© 2026 Rolando "Rollo" Alberti - All rights reserved
About Privacy Policy Cookie Policy