Particle.news
Download on the App Store

OpenAI's Astra Designated 'Critical' and Its Strongest Hacking Tools Will Be Restricted

The model can autonomously chain and exploit software flaws, forcing OpenAI to limit powerful cyber features to vetted defenders and add new controls.

Overview

  • OpenAI announced on Tuesday that Astra is the first of its models to meet the company’s Preparedness Framework 'Critical' cybersecurity threshold, a designation for models that can find and exploit previously unknown vulnerabilities without step‑by‑step human guidance.
  • In internal tests OpenAI says Astra scored 100 percent on an exploit-development benchmark called ExploitBench and discovered and used two previously unknown zero-day flaws as part of an exploit chain.
  • OpenAI paused parts of Astra’s work after July testing incidents, delayed the launch by several weeks, and says it has added stronger safeguards such as tightened sandboxes, a misalignment monitor, and training to make the model more likely to refuse unsafe cyber requests.
  • At release OpenAI will keep Astra’s most advanced cyber capabilities closed to the general public and grant early access only to a small set of vetted Daybreak partners, a move meant to let defenders harden systems without broadly enabling attackers.
  • The announcement has renewed debate over agentic model design and monitorability because architectural shifts like recurrent-depth may hide internal reasoning, and it has prompted industry calls for standardized pre-release testing, tighter operational controls, and seat‑belt rules for defensive use.