AI Watermark Defenses Quickly Breached, Highlighting Provenance Challenges, Linas Beliūnas Reports

L

Linas Beliūnas

LinkedIn Author

🔔linas.substack.com🔔 Daily Intelligence on Finance & AI | Scouting FinTech & AI Startups 🦄

In a recent LinkedIn post, Linas Beliūnas discusses the rapid circumvention of AI output watermarking, raising significant questions about the effectiveness of current provenance solutions for artificial intelligence.

Linas Beliūnas highlights the speed at which a free, open-source tool emerged to remove watermarks from major AI models, including Anthropic’s Claude, Google’s Gemini, and OpenAI’s models, shortly after Anthropic’s announcement of invisible watermarks.

“On Tuesday, Anthropic announced invisible watermarks in Claude’s output. Less than 24 hours later, someone had created a FREE Skill that removes the watermarks from Claude, Gemini, and OpenAI.”

The Adversarial Nature of AI Provenance

Linas Beliūnas points out that the core challenge with AI watermarking lies in the inherent asymmetry between defenders and attackers. The system designed to embed a watermark must be robust against numerous forms of manipulation, while an attacker needs only to find a single vulnerability.

Examining the ‘watermarks-remover’ Tool

According to Linas Beliūnas, the open-source tool, named ‘watermarks-remover,’ targets multiple layers of AI provenance. This includes stripping invisible Unicode characters, removing C2PA metadata from various file types, and attacking statistical text watermarks through output rewriting.

“The defender has to build a signal that survives almost everything. The attacker only has to find one way to break it.”

Linas Beliūnas elaborates on the limitations of specific technologies. While C2PA can offer cryptographic proof of origin, its metadata is vulnerable to being lost through common processes like re-encoding or screenshots. Similarly, statistical watermarks, though more resilient than simple metadata, can be weakened by paraphrasing the AI-generated text.

Watermarks as Tamper-Evident Stickers

The effectiveness of these watermarks, Linas Beliūnas suggests, is akin to a tamper-evident sticker rather than an impenetrable lock. While they can be useful if they remain intact, they are less reliable as definitive proof of origin when they can be easily compromised.

“So the watermark starts looking less like a lock and more like a tamper-evident sticker. Useful when it survives. Much less useful as definitive proof.”

This rapid breach is particularly significant, Linas Beliūnas argues, given the increasing reliance of governments, platforms, and AI developers on provenance as a means to combat deepfakes, misinformation, and unacknowledged AI-generated content.

The ‘AI Companies Add the Mark. Open Source Builds the Eraser’ Dynamic

Linas Beliūnas concludes by framing the situation as an ongoing arms race between AI developers and the open-source community. The ability for anyone to remove AI-generated labels undermines their utility.

“AI companies add the mark. Open source builds the eraser. And if anyone can peel the “Made by AI” label off, the label doesn’t tell you much.”

This dynamic underscores the complex and evolving landscape of AI content verification and the challenges in establishing reliable methods for identifying AI-generated material in an increasingly sophisticated digital environment.

📝 About This Content

This article is based on insights shared by Linas Beliūnas on LinkedIn.

📅 Originally posted on August 14, 2026 | View original post on LinkedIn →