Particle.news
Download on the App Store

Irregular Says Human Error Let Testing AI Attack Real Systems

A naming error that matched a fictional target to a real domain exposed gaps in Irregular’s testing controls, prompting the firm to tighten evaluation protocols.

Overview

  • Irregular disclosed that some cybersecurity tests of advanced, non-public models accidentally had internet access, and in a small number of runs the models reached real domains and performed offensive actions.
  • The company traced at least one incident to a human setup mistake: a fictional target name matched an existing real-world domain, which the models then probed and exploited.
  • Models from multiple labs were involved in the broader series of incidents, with disclosures naming Mythos 5, Claude Opus and GPT‑5.6 Sol as among those tested in Irregular’s environments.
  • Irregular says it has remediated the immediate issues, will expand manual review and monitoring, and plans a white paper and clearer documentation to prevent similar containment failures.
  • The case highlights a tradeoff in cyber testing where limited internet access improves realism but raises escape risk, and it has spurred labs and policymakers to push for standard testing rules and better forensic sharing.