How Claude marks AI-generated content:
When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.
Because the watermark is part of the text, it will travel with the text when it’s copied and pasted elsewhere, and may persist through some editing. Watermarking will be applied at the model level, which means it will be present no matter which Claude product or surface the text comes from.
Today Anthropic rolled out a new support page aimed at users in the European Union as part of signing the EU AI Act’s Article 50(2) Code of Practice on Transparency of AI-Generated Content. The new watermarking will apply to Claude models launched in the EU on or after August 2, 2026. Existing models are covered by a transition period but Anthropic says they’re working to add marking support to them too.
While I understand the concept of how this can work in theory, how it would actually work in practice is mind boggling. Even more so that this will apply even to API outputs from providers like AWS, Microsoft and Google.
How many characters will you have to use for something to be detectably Claude-marked? If I asked Claude to reply only with a single character, how can we be sure it’s marked? The answer is that we don’t know yet, and perhaps Anthropic has really just shipped an CYA page, especially as it pertains to marking older models.
Everyone seems to have an opinion on the effectiveness of “AI detectors” these days, but it will be interesting to see how well Anthropic’s works. If it’s any good, I expect other large economic powers will come knocking.
But I don’t see how the marking system they describe can be any good. Unless the Claude mark is complete gobbledygook (and they said it won’t be), there’s plausibly already text somewhere that would “match” this mark that wasn’t created by Claude. What’s the point of a mark if it’s more likely to cause confusion than no mark at all?
Imagine an author submits their original work to a publisher and they check it against an official Claude mark detector. Even a work with no use of AI whatsoever could may still not pass 100% because there’s nothing preventing someone from choosing the same words Claude might. And if doing so you choose to believe the author, why even check it in the first place?
So far it sounds like a system that solves for a very specific regulatory requirement, but no real problems. I’d anything, it might actually create more problems.
The case for open models just keeps getting better and better.