Anthropic will embed text watermarks into the output of future Claude models that support the feature, starting from their initial release. It will also add this to existing models launched before August 2, 2026, over the coming months. While the initiative begins as a response to the machine-readable labeling requirements in the EU market, Anthropic states that because it does not yet have a way to permanently separate functionality by region, it will apply the feature across all regions where the target models are offered. The announcement was made on August 14, 2026.

What changes here is that an entry point is provided for probabilistically investigating Claude's involvement. However, a positive detection result does not identify authorship or ownership. It does not prove violation of terms of service, fraud, or the truthfulness of content. A negative result likewise does not prove that text was human-written. Organizations that use the watermark need an operational approach that does not treat detection results as the sole basis for judgment.

Anthropic stated that it will make future models watermark-capable from their initial release, and will add this to existing Claude models launched before August 2, 2026, over the coming months. Therefore, not all Claude output is currently watermarked. It is necessary to separately examine the technical mechanism, the source of the evaluation figures, and the limits of detection.

AD

No Hidden Characters Added—Signal Remains in Word Choice

The method Anthropic describes does not add invisible character strings or metadata to text. When the model selects its next token, among candidates that would not compromise meaning, it slightly favors choices that align with a pseudo-random pattern derived from a secret key and the preceding context. The detector uses the same key to score how well the word order of the completed text matches that pattern.

SynthID-Text is a technique that alters sampling at generation time, rather than retraining the trained model. Provided the key and settings are kept secret, the detector can be made semi-public as an API. Conversely, this is not a technology that can detect output from providers who do not implement watermarking, or from openly distributed models running in decentralized environments. It also differs from a mechanism that identifies AI-generated text in general with a single classifier.

According to Anthropic, the watermark does not generate additional tokens, its impact on generation speed is negligible, and it does not change pricing. Anthropic states that the watermark itself does not include information identifying the user, organization, or conversation. This is explained not as an identifier that transmits content externally, but as a statistical trace left in the sequence of choices made during generation.

Meanwhile, Anthropic states that its internal testing found no impact on content, creativity, or readability. However, it has not disclosed the sample size, target models, or languages involved. The evaluation metrics and confidence intervals are also unknown, so there is still insufficient material for external parties to verify the quality impact specific to Claude.

The Evaluation of 20 Million Responses Is Not Claude's Own Measured Data

Anthropic describes the adopted method as "a version of SynthID-Text." The underlying peer-reviewed Nature paper was published on October 23, 2024, with authors from Google DeepMind and others evaluating it on Gemma 2B, Gemma 7B, and Mistral 7B-IT. This does not include an evaluation of Claude itself.

In Google's production-scale testing, approximately 20 million Gemini responses were randomly assigned to watermarked and non-watermarked groups. The watermarked group showed a 0.01% higher rate of high ratings and a 0.02% lower rate of low ratings, but neither difference was statistically significant. Additionally, in a human side-by-side comparison test of 3,000 ELI5 responses from Gemma 7B-IT, no significant differences were reported across five categories: grammar/coherence, relevance, correctness, helpfulness, and overall quality.

Latency figures in the paper also come with conditions attached. Under specific conditions running Gemma 7B-IT on 4 v5e TPUs, processing time per token increased from 15.527ms to 15.615ms with 30-layer Tournament sampling, a 0.57% increase. This is not a latency measurement for Claude. Since Claude's own configuration has not been published, the ngram length and number of Tournament sampling layers used in the Google paper cannot be assumed to match Claude's actual implementation values. The same applies to detection thresholds and accuracy figures broken down by language or model.

The result that SynthID-Text maintained quality at production scale is informative as a reference. However, it has not been confirmed that Claude will achieve the same quality, latency, and detection accuracy. Distinguishing between Anthropic's internal evaluation and the evaluation Google conducted on different model families is the starting point for post-deployment verification.

AD

Detection's Meaning Shifts With Short Text, Copyediting, and Translation

In Google's publicly available SynthID-Text implementation, detection results are returned in one of three states: watermarked, not watermarked, or uncertain. This same implementation also allows two thresholds to be configured, adjusting the relationship between false positive and false negative rates. Anthropic has not yet disclosed the return states or threshold design for the detection API it announced it would offer for Claude in the near future. At minimum, results from the Google implementation are not binary determinations, but probabilistic judgments that depend on the chosen thresholds and the characteristics of the text.

Longer texts make it easier to accumulate signal, since the model makes more selections among candidates. Conversely, the signal is weaker in short texts or factual statements where there is essentially only one correct answer. Strict copyediting and code also have little freedom of choice. However, signal can still be present in portions with expressive latitude, such as code comments.

Editing also has its boundaries. While minor corrections may leave the signal intact, comprehensive rewriting can erase it. The Nature paper states that paraphrasing by an LLM weakens the watermark, and Google's implementation documentation similarly explains that confidence can drop significantly after comprehensive rewriting or translation into another language.

Text translated by Claude falls within the scope of watermarking, since Claude chooses every word. When a human's text is only corrected for punctuation or grammar, few words change, and detection may fail. Furthermore, what detection reveals is a probability of involvement that does not distinguish whether Claude wrote the text or substantially edited it. If a positive result is used in content review or fraud investigation, conclusions cannot be drawn while that distinction remains unresolved.

The EU's Machine-Readable Labeling Requirement Extends to Claude Worldwide

Article 50(2) of the EU AI Act requires providers of AI systems to mark generated audio, image, video, or text content in a machine-readable format that enables detection as artificially generated or manipulated. The marking must be effective, interoperable, and robust and reliable to the extent technically feasible. Anthropic's announcement centers on this provider-side obligation.

The EU's Code of Practice is a voluntary framework supporting the implementation of Article 50. While participation in the Code itself is voluntary, the legal obligation under Article 50 is not. The Code was published on June 10, 2026, and by the end of July 2026, approximately 190 organizations had signed on. Section 1, aimed at providers, was signed by 82 organizations, including Anthropic and Google, as well as Meta, Microsoft, Mistral, and OpenAI.

Deployers have a separate role: when publishing AI-generated or manipulated text for the purpose of informing the public on matters of public interest, they must display a label. However, an exception applies when the content has undergone human review and editorial responsibility is assumed. Anthropic's text watermark cannot be treated as a mechanism that fulfills a deployer's labeling obligation. Who marks content at the generation stage and who applies labels when displaying it to the public are separate questions.

Existing systems are also subject to a transition period. The EU Commission's FAQ sets a deadline of December 2, 2026, for existing systems. Anthropic's plan to add the feature to existing Claude models over the coming months reflects, similarly to this transition period, a premise that already-deployed models will not be switched over all at once. However, the actual timing and scope per model will have to await future announcements.

AD

The Detection API and C2PA Should Be Treated as Separate Forms of Evidence

Anthropic announced that it will offer a watermark detection API in the near future. However, it has not disclosed the release date, pricing, or eligibility requirements. The relationship between thresholds and false positive/false negative rates, as well as Claude-specific accuracy, also remain unknown. Even once the API becomes available, users themselves will need to determine what length and type of text to examine at what threshold, together with an understanding of what the results actually mean.

Support for image files uses a different method from the statistical watermarking applied to text. For supported file types such as .png, .jpg, and .svg, Anthropic attaches a cryptographically signed C2PA Content Credential to the metadata, indicating that Claude generated or processed the file. C2PA is a standard that makes it easier to cryptographically verify a file's provenance; it is not a method that scores the sequence of words in text.

C2PA provenance information also does not determine whether the content of an image is true, and the embedded metadata itself can be removed. Understanding it as an indelible image watermark would be a misunderstanding of the technology's nature. Statistical text detection and provenance information are each only a piece of evidence.

Organizations that use Claude's watermark in practice need to decide in advance what additional verification steps to take for each of the positive, negative, and uncertain results. Even after the API's thresholds and Claude-specific accuracy figures are made public, detection results alone cannot serve as grounds for determining authorship or violations.