In the last few months, more AI content has meant a debate over who gets to label AI-generated content, how reliable those labels are, and whether people should even know when a machine has been involved in their writing.

In August, Anthropic announced that future Claude models would generate text with an invisible, machine-readable watermark. The company says the watermark does not add hidden characters or alter the visible quality of the text, but uses patterns in the model's word selection that can later be checked by a detector. The move is part of the company’s effort to comply with the EU AI Act.

Within hours of the announcement, Guillaume Meyer, a Los Angeles-based developer, began looking into how the technology worked. Meyer, who has spent 26 years in tech, is also the founder of Memo, an AI platform that analyses e-commerce brands’ advertising and performance data to help them improve their campaigns.

The project quickly gained traction, drawing thousands of stars and more than 100 contributors on Github. It is now called Watermarks Remover, an open-source tool that targets statistical watermarks created through a model’s choice of words. Instead of removing a hidden tag, the tool makes small changes to the wording and phrasing, while trying to keep the original meaning intact. The idea is to break the statistical pattern that a detector would use to identify the text as AI-generated.

Meyer says the tool is not simply designed to make AI-generated text “undetectable”. His objection is to the idea of an invisible, provider-controlled mark being used to determine whether AI was involved in creating someone's work.

“I am not against transparency,” Meyer told Decode. What he is wary of is the “invisible mark”.

“Having a private group of people label what people say without their knowledge or control can lead to very bad things. Most people don't even know that this exists. They have no way to control it and no way to audit it,” he said.

Decode spoke with Meyer about why he believes watermarks can produce false positives, how his tool tries to remove these marks while preserving meaning, and why he thinks AI companies may have incentives to watermark content beyond regulation.

Here are the edited excerpts from the interview.

What made you build an AI watermark remover after Anthropic's announcement?

The fun part of the story is that when I started reading about this, I started investigating the technology behind watermarking because I am a nerd. I've been in tech for 26 years, and I had no idea how it worked.

So I started reading about the technology and the research papers, mostly from Google, which was kind of a pioneer in this field. After a few hours of reading, I realised that some of the work I'd done in the past, mostly around open-source software and models, could potentially be useful to circumvent these watermarks.

It was completely out of curiosity. I started building something just to see if it worked. And it looked like it was going to work. It took me something like five hours to ship the first prototype.

At the beginning, I had zero idea whether it was right or wrong. It was purely a technical research project. But as I started seeing the reactions, I realised that many people had strong opinions about it.

What changed your view from seeing this as a technical experiment to questioning the use of AI watermarks more broadly?

I am actually all for content attribution. I have no problem with content attribution. What I don't like is the watermarking technique, for several reasons.

The first is philosophical. From a historical perspective, having a private group of people label what people say without their knowledge or control can lead to very bad things. Most people don't even know that this exists. They have no way to control it and no way to audit it.

Then there is a practical problem. Even Anthropic says in its paper that being detected doesn't mean it is necessarily true, and not being detected doesn't mean it is necessarily wrong. So what is the practical value?

And there is another problem. I am not a native English speaker and I use Grammarly all day long to proofread my writing. Grammarly uses LLMs. So if those systems start watermarking text, potentially my writing could get a watermark simply because I am checking my grammar.

The technique also treats authorship almost as zero or one. You can ask Claude to write an entire book, or you can ask it to change a few lines, reorganise your chapters or proofread something, and you can potentially be flagged in the same way.

That can create false positives. Imagine you're applying for a job and an organisation uses an AI detector. Your application gets flagged as AI-generated and you are rejected. You may never even know why, and you have no way to prove that the result was wrong.

Before getting into how your tool works, can you explain what an AI watermark actually looks like inside a piece of text?

The older watermarking techniques were quite simple. They would insert what we call control characters — invisible characters in the text. That was very easy to evade because you could copy and paste the text into a plain-text editor and copy it back, and those characters would disappear.

The newer watermarking technique is more clever. Instead of adding a visible mark, it subtly changes the words Claude chooses. For example, if the sentence is “The protest was very crowded,” the system might choose “extremely crowded” instead. Both mean almost the same thing, so you would not notice anything unusual.

It does this with word choices throughout the text, following a hidden pattern based on a secret key.

If you read the output, there is no obvious way to tell that anything has been changed. But the detector has that secret key. It looks at the pattern of word choices across the text and checks whether they match the pattern expected from Claude. From this, it can estimate whether the text was likely written by Claude.

But the detector relying on a secret key that only the provider holds is also a problem for me. If private companies are using private keys for a detector, editors or other independent organisations cannot actually check what's happening.

So I think there is a problem even from the regulatory perspective. It creates bad incentives that can lead to bad behaviour.

How does your tool try to remove that watermark while keeping the original text as intact as possible?

The simple approach would be to take a model that isn't watermarking text, give it the output of Claude or Gemini and ask it to rewrite it — find synonyms, slightly change the structure, reorganise things, make it sound more human. You would end up with a text that says the same thing but is written differently, and the original statistical pattern should be randomised.

That's the naive approach.

What we have built is slightly more complex because we want the text to be as close as possible to the original while still not containing the watermark. It's not a deterministic process. We make small edits, check what the text says, see whether we've changed the original meaning, and repeat that several times. We also generate multiple versions of the same text and compare them.

We are generating multiple variants and checking them at each step. Technically, we are calculating what we call a semantic distance. It gives you a number representing how much the meaning has drifted.

So when people tell me, "You are just rewriting the content," no, we're not just rewriting the content. It's a tricky process.

And this is a long-term project. It blew up because I published it quickly, but it will require a lot of work from the community. We already have contributors helping. When Anthropic publishes its detector, we'll have to make adjustments. The watermarking technology will progress, and we will have to progress at the same time.

If you don't think invisible watermarks are the right solution, what would you like to see instead?

I am completely fine with disclosure. If you're sharing something that is 100% or 98% generated with AI, you can disclose that. I do it myself. If you're using AI to generate pictures or videos of real people, there should obviously be a visible watermark or some form of disclosure.

I actually hate the idea of invisible watermarks. If I am writing a piece of text and using AI to help me, I don't think that an invisible mark inserted into my words is the right solution. There is no perfect solution. I understand that, and I am open to good arguments from people who support watermarking.

But I do think there are incentives for AI companies for introducing watermarking that go beyond regulation.

When you are training a new model, you're reading everything that's online. But now a huge amount of content online is AI-generated, and that isn't necessarily useful for model providers because they want to avoid training on their own generated content. Watermarking can help them identify content they generated and potentially exclude it from future training data.

The second incentive is model distillation. If another company takes your model and uses its outputs to train another model, watermarks could potentially help you identify that. That makes it harder and more expensive for other companies to distill your model.

I think those are significant incentives, and nobody really talks about them.

That's why I am happy the project has started a broader conversation. It is open source and MIT-licensed, and it will remain that way. I may find ways to monetise something around it because I am an entrepreneur, but the project itself will remain open source.

About the author
Hera Rizwan
Hera Rizwan

Hera Rizwan is a correspondent with Decode. She covers AI, technology, and accountability, with a focus on how digital systems shape welfare, governance, and public life in India. Her work examines the real-world impact of emerging technologies, from biometric systems and surveillance tools to platform-driven scams and digital policy. She has reported extensively on cybercrime, AI in welfare delivery, and the intersection of tech and democracy. She is a Pulitzer Grantee and won Ramnath Goenka Award for Investigative Reporting in 2022.