In brief: Tools circulating that claim to remove Claude’s text watermark can demonstrably only delete metadata, not the actual word-choice-based watermark—yet they bring their own supply-chain risks via dynamically loaded third-party archives.
Just a few days after the introduction of the text watermark for Claude, several tools promising its removal are already circulating. Independent testing, however, shows that they cannot actually attack the real, word-choice-based watermark at all—but they do bring their own risks to the software supply chain.
Since August 2, 2025, Anthropic has been embedding an invisible watermark woven into the word choice of text generated by newer Claude models. Supported file formats additionally receive signed C2PA metadata. The trigger is Article 50 of the European AI Act, which has been enforceable since the same date and provides for fines of up to 15 million euros or 3 percent of global annual turnover for violations. Just a few days later, several tools promising to remove this watermark are already circulating—including the open-source, MIT-licensed project watermarks-remover by developer Guillaume Meyer, other GitHub projects, newly registered websites, and a corresponding offering from the commercial provider StealthGPT, which specializes in evading AI detection software.
According to the trade publication BleepingComputer, only two of the advertised functions can actually be verified: removing hidden, invisible characters in the text and deleting C2PA, EXIF, and XMP metadata from files—effects that already occur anyway through re-saving, a format conversion, or a screenshot. However, the actual watermark is not contained in hidden characters but in the model’s specific word choice, and according to current knowledge could only be removed through comprehensive rephrasing using a second model. Meyer himself publicly admitted that his tool so far only removes metadata; actual removal of the word-choice-based watermark is not currently available. Security researcher Pasquale Pillitteri also analyzed the source code of several projects and found that one widely used tool leaves the most common technique for hiding payload data in text completely untouched.
For CISOs, the limited technical effectiveness of these tools is less relevant than their practical integration. Anthropic itself points out that a detected watermark merely means that a text has been processed by Claude—not necessarily fully authored by Claude, since even a grammar check, a translation, or a summary performed by the model triggers the marking. The company announces it will support external detection capabilities in the future and publish further technical documentation.
BleepingComputer sees the real danger in the practical use of such tools: the watermarks-remover project can be integrated directly into existing AI agent workflows as a so-called agent skill and, via an optional evaluation function, additionally downloads an archive of roughly 220 megabytes from a third-party repository—a classic vector for supply-chain risks. The magazine emphasizes that it has not itself tested any of the tools and recommends the same caution as with any other unverified code from the internet. While the current projects at least offer openly viewable source code, BleepingComputer warns that an upcoming wave of similar tools, given the demand already demonstrated, could be significantly less transparent. Security officers are therefore advised to scrutinize agent-skill integrations and third-party download functions in AI workflows particularly critically before deploying them in production environments.
Source: www.it-daily.net · Published August 14, 2026
Lumi AI News — AI-assisted curation pursuant to Art. 50 EU AI Act. Paraphrasing and classification by Lumi News Pipeline v1.8.3.