Overview
- Anthropic disclosed on Tuesday that it has started embedding an imperceptible, machine‑readable watermark into text produced by Claude models released on or after August 2 and that it will extend marking to older models.
- The mark is a statistical fingerprint added at the token‑selection level: the model biases choices among near‑equally likely words so a pattern emerges that a holder of the key can detect.
- Anthropic says the watermark can flag text the model only edited or proofread as well as text it fully generated and that it does not identify users, but the company has not yet published the detection tool or empirical error rates.
- Within hours of the announcement independent developers published open‑source 'watermark removers' and rewrite techniques—using paraphrase, translation, condensation, or other models—to weaken or erase the statistical signal.
- The change is driven by the EU AI Act’s transparency requirements and raises practical risks for students, job applicants, authors and coders because the mark is probabilistic, can be lost after heavy editing, and may produce misleading attributions.