Overview
- Anthropic published technical notes this week explaining that Claude models launched after early August will include an invisible watermark created by biasing low‑stakes token choices according to a secret key.
- The method changes the model’s randomness at inference so nothing is added to the text and no hidden characters or user identifiers are embedded.
- Anthropic says the watermark is strongest in long, free‑wording outputs, is weaker or absent in single‑answer factual replies and much code, and can be broken by extensive rewriting.
- The company plans to offer a watermark detection API and will attach C2PA cryptographic provenance to supported files, but it has not published a public detector, detection thresholds, or empirical error rates.
- The approach reuse concepts from the 2024 SynthID work and is being rolled out globally to ensure EU compliance, a move that raises concerns about verification power concentrating with providers and a potential arms race over removal and circumvention.