Ads_970x250

How Anthropic Plans to Watermark Text Generated by Claude

Anthropic will mark text generated by future Claude models with a statistical signature that readers cannot see, as AI companies adjust to new European Union transparency rules.

Topics

  • Anthropic is embedding an invisible watermark into text generated by future Claude models in what could establish whether Claude had a hand in writing a passage without adding labels, hidden characters or metadata to the words themselves.

    The company detailed the system on Friday, 14 August, saying the watermark will be built into the choices Claude makes as it generates text. Anthropic plans to use it globally rather than only in Europe, at least initially, as it works to comply with transparency requirements under the European Union’s AI Act.

    There will be nothing for a reader to spot. Instead, the system leaves a statistical pattern in the words Claude chooses.

    Statistical Pattern

    Large language models (LLMs) produce text by repeatedly choosing the next token from a set of possible candidates. Anthropic’s watermark changes the way Claude makes some of those choices when several answers would work equally well. Someone with the corresponding detection key can then examine a longer passage and calculate how closely its sequence of choices matches the pattern expected from Claude.

    Anthropic is using a version of SynthID Text, a watermarking technique developed by Google DeepMind and described in a 2024 paper in Nature. The method changes the sampling process used to select tokens rather than inserting extra characters into the finished text. Google tested SynthID on live Gemini traffic and reported no statistically significant difference in user ratings between watermarked and unwatermarked responses.

    That does not make the watermark a definitive AI detector.

    Anthropic says its system can answer a narrower question: how likely is it that Claude was involved in producing the text?

    It cannot establish that a passage was written entirely by Claude, distinguish writing from heavy editing, or detect text produced by another AI system using a different watermark. Short passages are also harder to assess because they contain fewer word choices from which a reliable statistical pattern can emerge.

    Factual writing creates another problem. When there is only one sensible answer, there is little room for the model to vary its choice without risking an error. The same is true of computer code. Anthropic said lightly proofreading a piece written by a person may leave too little Claude-generated material for the watermark to be detected.

    Editing can weaken it as well. Anthropic expects the signal to survive some light changes, but says a complete rewrite can remove it. Google DeepMind has made much the same point about SynthID, saying detection works better on longer, varied writing and can lose confidence after substantial rewriting or translation. DeepMind has previously described the technology as useful but “not a silver bullet” for identifying AI-generated material.

    Legal Requirements

    The push is being driven by Europe.

    Article 50 of the EU AI Act requires providers of generative AI systems to make synthetic text, audio, images and video machine-readable and detectable as artificially generated or manipulated. Those transparency provisions became applicable on 2 August. The law says the technical measures should be effective, interoperable, robust and reliable as far as technically feasible.

    That requirement is separate from rules governing what people publishing AI-generated material must disclose to readers. For text dealing with matters of public interest, the AI Act requires disclosure when AI generates or manipulates the material, but provides an exception when the content has gone through human review or editorial control and a person or organization assumes editorial responsibility for its publication.

    Anthropic signed the EU’s voluntary Code of Practice on Transparency of AI-Generated Content, which sets out a route for companies to meet the legal requirements. About 190 organizations had signed by the end of July. Providers signing the relevant section include OpenAI, Google, Meta, Microsoft, Mistral and Cohere as well as Anthropic.

    Watermark-detection API

    Anthropic said it will apply the watermark globally at launch because it does not yet have a reliable way to restrict the system by geography. Models released before 2 August are covered by a transition period, and the company said it plans to add watermarking to them over the coming months.

    For users, checking the mark will require another piece that is not yet available. Anthropic said it is preparing a watermark-detection API, but has not given a release date. The watermark itself contains no information identifying the user, organization or Claude conversation, according to the company.

    Files will be handled differently. Supported formats such as PNG, JPG and SVG will carry cryptographically signed C2PA content credentials in their metadata indicating that Claude made or processed the file. Unlike the text watermark, that information sits in the file rather than in the statistical pattern of the output.

    So the watermark is evidence of involvement, not authorship. That distinction could become important as AI moves deeper into writing, editing and research workflows. A passage carrying Claude’s mark may have been generated from scratch, substantially rewritten or simply passed through the model. The watermark, by itself, cannot tell which happened.

    Topics

    More Like This

    You must to post a comment.

    First time here? : Comment on articles and get access to many more articles.