How the Anthropic Watermark Really Works in Claude

How the Anthropic Watermark Really Works in Claude

TL;DR Summary:

Word-Choice Watermark: Anthropic says Claude’s watermark is not hidden characters or style tricks, but a change in how the model chooses each next word during generation.

Hard to Strip: Light editing usually will not remove it, but a full rewrite can erase the pattern, while short or fact-heavy text is harder to detect reliably.

Built on SynthID: The system is based on Google DeepMind’s SynthID-Text and a 2022 watermarking idea, with Anthropic planning a detection API and EU AI Act support.

MirrorMark Angle: A newer method called MirrorMark may be even more resistant to editing, but Anthropic has not confirmed Claude uses it.

How does Anthropic’s watermark on Claude’s text actually work? Anthropic explained the method on August 15, 2026, and confirmed that the Anthropic watermark does not rely on hidden characters, em dashes, or writing style. Instead, it changes how Claude picks words during text generation. This matters now because it settles months of guessing about how AI detection tools work and shows what does and does not defeat them.

What The Anthropic Watermark Actually Changes

Claude generates text by picking likely next words, but with some randomness built in. The Anthropic watermark does not add anything to the text. It changes the source of that randomness. Instead of using a random number generator, the system uses a watermark key combined with the words that came before to decide which word Claude picks next. Anyone with the key can check whether a sequence of words matches the pattern that key would produce. Anthropic says this has no effect on quality and does not push Claude toward unusual word choices. For writers and editors working with AI-generated drafts, this distinction matters: watermarking operates at the level of word-choice statistics, not tone or phrasing, so it has nothing to do with preserving a consistent authorial voice across edits. Tools built for that purpose, like WordHero, address a separate problem entirely, helping maintain steady, on-brand wording through multiple rounds of rewriting rather than altering the invisible word-level patterns a watermark relies on.

The Anthropic Watermark Builds On SynthID-Text

Anthropic confirmed its approach is a version of SynthID-Text, a method Google DeepMind published in a Nature paper in 2024. The idea traces back further, to a 2022 proposal by Scott Aaronson. Anthropic’s version is not identical to SynthID. Two years of development separate the original method from what Claude uses today, though Anthropic has not detailed every difference.

How The Anthropic Watermark Can Be Defeated

Anthropic states directly that light editing likely will not remove the watermark. A complete rewrite that replaces every word will remove it. Anthropic also notes that in that case, you could argue the text no longer counts as AI-generated at all. Detection also weakens in specific conditions. Short samples reduce accuracy, since there are fewer word choices to check against the key. Factual content weakens detection too, because fewer word options exist when facts constrain the choices. Simple proofreading, where only grammar and punctuation change, leaves too few altered words for the watermark to register reliably.

What Comes Next For The Anthropic Watermark

Anthropic plans to release a detection API so others can check text against the watermark key. Future Claude models will include watermarking as part of compliance with the EU AI Act. For non-text image files, including JPG, PNG, and SVG formats, Anthropic will use C2PA metadata instead, a different labeling standard built for images rather than text. Anthropic says the watermarking process adds no extra token cost and has a trivial impact on generation speed.

A Newer Method Called MirrorMark Goes Further

A separate 2026 method called MirrorMark extends the SynthID approach. It spreads watermark information across a document using a Context-Anchored Balanced Scheduler, or CABS, which decides where each part of the watermark sits based on surrounding words. Unlike SynthID’s zero-bit design, which only signals watermark or no watermark, MirrorMark can encode multiple bits of information across the text. This makes it more resistant to editing. Anthropic has not confirmed that Claude uses MirrorMark specifically, and that connection remains unconfirmed.

Watch for Anthropic’s detection API release, since that will let you test text directly against the watermark key. Remember that a full rewrite can remove the watermark, while light edits generally will not. If you write or edit AI-generated text professionally, understanding these limits matters more than tracking any single detection tool. If you’re producing AI-assisted content regularly and want consistent, on-brand phrasing that still holds up to editing without triggering awkward rewrites, tools like WordHero can help maintain a steady voice across drafts as you navigate these detection and watermarking realities.


Scroll to Top