Anthropic signed the EU AI Act's Code of Practice on Transparency of AI-Generated Content, and started marking text produced by Claude with an invisible statistical watermark. Within days, the same event got read two completely different ways on my feed.
{%embed https://dev.to/sylwia-lask/the-end-of-undetectable-ai-text-claudes-new-watermark-explained-45g2 %}
She actually went and read the official documentation and reported what Anthropic really says: the watermark doesn't prove a text was entirely AI-generated – you can write something yourself, have Claude correct it or translate it, and it can still carry the mark. The reverse holds too: no watermark detected doesn't prove a human wrote everything.
The second, on Medium, called the announcement a "nuclear bomb" for the AI-generation community, and claimed "you are effectively carrying a digital scarlet letter." No citation of the actual documentation. Just outrage, with the appropriate amount of dramatic punctuation.
Between the two, I left a comment under Sylwia's article. Written in French, translated by ChatGPT, signed off with a question: how do you classify a text thought through, written, and reviewed by a human, but whose final English phrasing came out of a model? It wasn't a rhetorical question. I still don't have an answer.
The Medium article mixes together, without ever saying so, three mechanisms that have almost nothing to do with each other.
Anthropic's watermark. A statistical signal baked into the model, documented, which openly acknowledges its own limits: it can disappear after enough editing or translation, and its presence says nothing about who came up with the idea in the first place.
Third-party consumer detectors. ZeroGPT and its relatives. Tools that have existed for years, unrelated to Anthropic, whose reliability has never been seriously demonstrated at scale.
Platform decisions. A badge displayed on Medium, a curation algorithm that penalizes flagged content. These are editorial choices specific to each platform, not a mechanical consequence of the watermark.
The Medium article writes as though the first link automatically triggers the next two: Anthropic marks, so detectors will catch everything, so platforms will punish. The reasoning collapses the moment you separate the links – and for good reason: detecting that a model was involved in producing a text is not the same as determining that the text was "written by AI." That conflation hides a deeper one, which public debate almost systematically ignores.
A text assisted by AI stays under human editorial control from start to finish: the idea, the angle, the structure, and the final decision to publish belong to someone, no matter how many back-and-forths with a model happened along the way to correct, rephrase, translate, or pressure-test a line of reasoning. A text generated by AI comes out of a prompt, without that upstream control – the idea itself came from the human, but the generated text, as it stands, belongs to the model. A text produced by AI at industrial scale is something else again: an automated publishing pipeline, without anything resembling editorial oversight.
What separates the three, then, isn't how much the model intervened – it's who kept their hand on the decisions.
These three cases carry entirely different editorial responsibility. Treating them as a single category ("AI content") means judging a text on a binary criterion where the reality is a full spectrum of nuance – which is exactly what an undifferentiated "processed by AI" badge does.
I ran an article I wrote in April 2021 – before any consumer-facing LLM existed – through ZeroGPT. A test about a year ago gave it a 97% AI probability. A recent test, on the exact same text, with the exact same tool, gives 8.6%.
Same tool. Same text, down to the punctuation. Two incompatible scores, a year apart.
Digging into the flagged passages in the second test, a pattern emerges. These aren't random sentences: they're consistently the most neutral, most pedagogical, most well-structured passages in the piece – a definition of what a web server is, an explanation of CentOS Stream, a step-by-step automated update procedure. The passages where my voice actually comes through – the self-deprecation, the verbal tics, the Neapolitan moka pot bought at a flea market – are never flagged.
One test doesn't prove what a detection model measures in general. But this one strongly suggests the tool reacts to stylistic neutrality and structural regularity, not to a text's actual origin. Which is precisely the problem: those characteristics existed in human writing long before LLMs did. The detector is chasing a style whose origin is human, using the machine's imitation of it as the reference point. The reasoning eats its own tail.
Since the whole point of this article is how hard it is to judge a text by its origin rather than its content, it's worth being transparent about how this one was produced.
