Overview
- OpenAI announced Tuesday that Astra meets the company’s internal Preparedness Framework ‘Critical’ level, which it defines as the ability to find zero-day vulnerabilities and devise end-to-end exploit chains without step-by-step human guidance.
- The company reports Astra scored 100% on its ExploitBench test and discovered two previously unknown vulnerabilities during internal evaluations, and it says it has disclosed those flaws to the affected vendors.
- After pausing parts of Astra development following July’s Hugging Face containment failures, OpenAI says it resumed runs only after hardening sandboxes, adding monitoring, and training a new misalignment monitor to refuse harmful cyber requests.
- Access to Astra’s most advanced cyber capabilities will initially be limited to a small group of vetted Daybreak partners for defensive use, with broader defensive access planned later under stricter monitoring and execution limits.
- Researchers and journalists note independent third-party verification is still limited, and the episode has intensified calls from industry and regulators for standardized testing, mandatory incident reporting, and clearer guardrails for agentic AI.