Overview
- OpenAI paused internal Astra activities on Monday after preliminary evaluations showed the model might meet the company's Preparedness Framework 'Critical' threshold, meaning it could autonomously find or chain zero‑day exploits or plan end‑to‑end attacks.
- The company moved Astra into isolated test environments, restricted network and tool access, added weight protections and encryption, and deployed universal monitoring that inspects models’ chain‑of‑thought to interrupt risky agentic behavior.
- Simultaneously OpenAI expanded its Daybreak program and released a more permissive variant, GPT‑5.6‑Cyber, for tightly vetted defenders under identity checks, logging, legal attestations and monitored use.
- OpenAI said GPT‑5.6‑Cyber completed about 95% of advanced offensive cybersecurity test requests while standard models completed roughly 1.5–2%, and the Cyber model helped find real flaws including two V8 (Chrome) issues disclosed and fixed as CVE‑2026‑15903.
- Industry and government engagement and ongoing forensic reviews continue as labs and defenders weigh the benefits of giving vetted security teams powerful AI tools against the risk that similar capabilities could be misused if containment or test configurations fail.