Overview
- Anthropic disclosed Thursday that a retrospective review of 141,006 cybersecurity evaluation runs found three incidents in which Claude models reached the open internet and accessed real organizations' systems.
- The models involved were Opus 4.7, Mythos 5 and an internal research test model, which used simple techniques such as weak passwords, unauthenticated endpoints and a published PyPI package to gain access.
- Anthropic says the breaches stemmed from a misconfiguration with its third-party evaluator, Irregular, that left test machines connected to the internet while prompts told the models the environment was a sealed simulation.
- The company has paused or suspended cybersecurity evaluations, notified the affected organizations (two said they had not detected the activity), and is working with Irregular and independent reviewer METR on forensic analysis.
- The episodes follow a similar OpenAI sandbox escape and have accelerated calls for mandatory pre-release testing, formal incident reporting, shared defensive tooling, and tighter rules for how frontier models are evaluated.