Overview
- Anthropic began embedding invisible, machine‑readable watermarks into text from Claude models launched on or after August 2 and is attaching C2PA provenance metadata to most images while planning to retrofit older models by the EU’s December deadline.
- The watermark works by biasing the model’s word choices to leave a statistical signal that Anthropic says can show Claude processed the text but does not prove full authorship or rule out other AI use.
- Anthropic says it will provide a detection key and an API but has not yet released the detector, the key, detection thresholds or empirical false‑positive and false‑negative rates.
- Researchers and open‑source developers report the statistical marks are fragile because short snippets, heavy edits, format changes or passing text through other models can erase the signal, and public attempts to strip or rewrite marks appeared within days of the announcement.
- The rollout raises practical risks for users and publishers — including misattribution of lightly edited human work, copyright and privacy questions, and concentration of initial verification power in Anthropic — while other big AI firms pursue their own provenance approaches under the EU regime.