Particle.news
Download on the App Store

Anthropic to Embed Secret‑Key Watermarks in Claude’s Text Outputs

Built to comply with the EU AI Act, the system embeds a machine‑readable signal with detection access and error rates currently undisclosed.

Overview

  • Anthropic published technical notes this week explaining that Claude models launched after early August will include an invisible watermark created by biasing low‑stakes token choices according to a secret key.
  • The method changes the model’s randomness at inference so nothing is added to the text and no hidden characters or user identifiers are embedded.
  • Anthropic says the watermark is strongest in long, free‑wording outputs, is weaker or absent in single‑answer factual replies and much code, and can be broken by extensive rewriting.
  • The company plans to offer a watermark detection API and will attach C2PA cryptographic provenance to supported files, but it has not published a public detector, detection thresholds, or empirical error rates.
  • The approach reuse concepts from the 2024 SynthID work and is being rolled out globally to ensure EU compliance, a move that raises concerns about verification power concentrating with providers and a potential arms race over removal and circumvention.