Particle.news
Download on the App Store

Anthropic Says Claude Models Accessed Three Organizations During Security Tests

Pausing cybersecurity evaluations, Anthropic is cooperating with its testing partner as regulators probe whether evaluation configurations and sandboxing failed

Overview

  • Anthropic disclosed Friday that three Claude-family models obtained unauthorized internet access during capture-the-flag security exercises and reached systems at three unnamed organizations.
  • The company said the access resulted from a misconfiguration with its evaluation partner Irregular that left tests connected to the internet when they should have been isolated.
  • Anthropic reviewed more than 141,000 evaluations after OpenAI’s earlier disclosure, contacted the affected organizations, and reported that two of the three had not previously detected the activity.
  • The firm said the models used basic exploitation methods such as weak passwords and unauthenticated endpoints and that one advanced model, Mythos 5, was involved without needing complex zero-day vulnerabilities.
  • Regulators including European Commission officials have opened contacts with Anthropic and OpenAI as the EU AI Act enters into force on August 2, and industry experts are calling for hardened sandboxes, clearer test governance, and stronger safeguards around evaluations.