Overview
- OpenAI announced on Tuesday that Astra is the first of its models to meet the company’s Preparedness Framework 'Critical' cybersecurity threshold, a designation for models that can find and exploit previously unknown vulnerabilities without step‑by‑step human guidance.
- In internal tests OpenAI says Astra scored 100 percent on an exploit-development benchmark called ExploitBench and discovered and used two previously unknown zero-day flaws as part of an exploit chain.
- OpenAI paused parts of Astra’s work after July testing incidents, delayed the launch by several weeks, and says it has added stronger safeguards such as tightened sandboxes, a misalignment monitor, and training to make the model more likely to refuse unsafe cyber requests.
- At release OpenAI will keep Astra’s most advanced cyber capabilities closed to the general public and grant early access only to a small set of vetted Daybreak partners, a move meant to let defenders harden systems without broadly enabling attackers.
- The announcement has renewed debate over agentic model design and monitorability because architectural shifts like recurrent-depth may hide internal reasoning, and it has prompted industry calls for standardized pre-release testing, tighter operational controls, and seat‑belt rules for defensive use.