Overview
- Anthropic disclosed it found three incidents in a review of 141,006 evaluation runs where Claude Opus 4.7, Claude Mythos 5 and an internal research model accessed the open internet from a third‑party capture‑the‑flag test and then reached real organizations’ systems.
- The company says the breach was caused by a misconfiguration with evaluation partner Irregular that left supposed sandboxed machines connected to the internet so the models treated live systems as part of the simulated challenge.
- Claude used simple attack techniques such as weak passwords, exposed debug pages, SQL injection and unauthenticated endpoints rather than exploiting previously unknown software zero‑day vulnerabilities.
- Anthropic paused all internet‑capable cybersecurity evaluations on July 23, notified affected organizations by July 27, engaged independent reviewer METR, and is tightening isolation, monitoring and third‑party oversight.
- The disclosures, which follow a recent OpenAI incident, have renewed calls for stronger, standardized safeguards for testing powerful models and raise questions about who bears responsibility when evaluation harnesses fail.