Anthropic Explains Claude’s Watermarking Process in New Blog Post
Anthropic released a blog post on Friday addressing fundamental questions about the watermarking of text generated by its chatbot, Claude. This move comes as part of the company’s commitment to comply with the EU AI Act’s Transparency Code, which mandates that AI companies implement systems to identify AI-generated content. The announcement has sparked debate among Claude users, with some voicing concerns about potential deception associated with the watermarking.
On platforms like Reddit, reactions have varied; some users labeled the watermarking as a conspiracy against Claude’s users, while others accused detractors of wishing to mislead. Business Insider noted that dozens of users on X have reportedly canceled their Claude subscriptions in response to this announcement.
How Watermarking Works
Anthropic’s blog details the watermarking concept, explaining that Claude generates a unique pattern in its text during “low-stakes choices,” such as selecting between descriptive words. This pattern remains “undetectable to the reader” but can be identified by anyone with the necessary key. The company emphasized that watermarking will not affect the quality of Claude’s output, stating, “To a reader, a watermarked response is indistinguishable from an unwatermarked one.”
Specifically, Anthropic plans to utilize the SynthID-Text approach developed by Google DeepMind in 2024 and aims to release a watermark detection API. The company clarified that watermarking differs from AI detection methods employed by companies like Pangram, which identify writing “tells” as evidence of AI use.
Editing and Code Implications
Regarding the possibility of rewriting text to remove the watermark, Anthropic indicated that while light editing may not eliminate the watermark, a complete rewrite would. However, the company acknowledged that such a case raises questions about whether the text could still be classified as AI-generated.
For text only proofread or lightly edited by Claude, the detectability of the watermark depends on the text’s length and the extent of the edits made. If the editing is minor, “nearly all the words” may be attributed to the human author, minimizing the watermark’s presence.
In terms of code generation, Anthropic noted that the watermark would likely be less prominent since the model must produce functional code with limited discretion for alternative word choices. However, the company acknowledged that in specific areas, such as code comments, where arbitrary terms may exist, watermarking could still occur.
Furthermore, Anthropic indicated that Claude is not alone in this watermarking initiative, noting that “other major model developers have signed the same Code of Practice and will be implementing their own watermarks.”


