Overview
- Anthropic began applying a machine‑readable watermark to Claude outputs for models launched on or after Aug. 2, 2026, and plans to retrofit older models by December.
- The watermark works by biasing token selection so word choices carry a subtle, statistical signal that detectors can spot even after copy‑paste and some editing.
- Anthropic has not published the detection API keys, thresholds, or empirical false‑positive and false‑negative rates, leaving uncertainty about how reliably the mark can be used in discipline or legal settings.
- Users and developers have pushed back by cancelling subscriptions and releasing tools and paraphrasing methods that can weaken or remove the watermark, showing it is not foolproof for short passages or heavy edits.
- Experts note the move may exceed the EU’s Article 50(2) exemption for light editing, raises questions about attribution and access for schools and companies, and follows a technical lineage that includes Google’s SynthID and earlier academic watermark proposals.