Anthropic launches watermark verification API for Claude-generated text, open to regulators, media and researchers
Anthropic launched an API that lets approved organizations (regulators, law enforcement, media, fact-checkers, independent researchers, educational organizations, EU civil society groups, and compliance-needing enterprises) check whether a piece of text contains Claude's invisible digital watermark. The watermarking itself (built on Google's SynthID text method, with tweaked word-selection randomness) was already required in Claude output since August 2025 under the EU AI Act; what's new is the verification/detection access being opened up.
Entities: Anthropic, Claude, Google, SynthID, Pangram, EU AI Act
0 primary
What happened
Anthropic has opened up a verification API that lets approved regulators, law enforcement, media, fact-checkers, researchers, EU civil society groups and compliance-focused enterprises check whether text was generated by Claude. The underlying watermark itself is not new: it is based on Google's SynthID method and has been embedded in Claude output since August 2025 to meet EU AI Act transparency rules. What changed today is that outside parties can now request access to verify that watermark, rather than Anthropic being the only one able to check.
Why it matters
This gives regulators, journalists and researchers a tool to check AI provenance without relying solely on third-party detectors like Pangram, which is useful for EU AI Act compliance and misinformation checks. But access is gated by Anthropic's approval process, so the company still controls who can verify what, and the practical impact depends entirely on how open or restrictive that gate turns out to be. For enterprises, this could also cut both ways: verifiable watermarks may satisfy compliance teams but could also expose AI use in contexts where contracts ban AI-generated content.
What is noise
The claim that this is "far more reliable" than existing detectors like Pangram is unquantified. No detection accuracy, false-positive rate, or false-negative rate has been published. The assertion that watermarking has no effect on content quality is disputed by critics and is essentially untestable from the outside. The phrase "may persist through some editing" is a hedge, not a robustness claim, and should not be read as durable protection against paraphrasing or editing.
Watch next
- 01Whether Anthropic publishes actual detection accuracy and false-positive rates for the watermark, rather than just claiming reliability
- 02How the approval process works in practice: how long applications take, who gets rejected, and whether access stays limited to named categories or expands to the public
- 03Whether independent researchers or outlets like Pangram test and confirm the watermark's persistence through common edits (paraphrasing, translation, partial rewrites)
Coverage
1 storyMore regulation signals
Full feed →- New York State legislature passes one-year moratorium on new large data centers5 Jun 202692
- Cloudflare mandates AI companies to separate web crawlers for search and training1 Jul 202690
- FERC mandates fast lane for data center interconnections to the grid18 Jun 202682
- Police officer investigated for using AI to create evidence in multiple cases13 Jun 202682