Particle.news
Download on the App Store

OpenAI Pauses Astra Development After Tests Flag 'Critical' Cyber Risk

The company is moving work into locked sandboxes so outside experts and officials can check whether the model can autonomously find and weaponize zero‑day flaws before any wider release.

Overview

  • OpenAI said on Friday that preliminary internal evaluations showed Astra may meet its highest cybersecurity threshold, meaning the model could autonomously identify and exploit severe software vulnerabilities.
  • The company has paused internal Astra activities that do not meet strengthened controls and moved further work into isolated, network‑restricted sandboxes with tighter protections for model weights and real‑time monitoring.
  • OpenAI confirmed Astra was not involved in the July breach of Hugging Face and said the pause will let government agencies, independent safety groups and selected third‑party testers run guided evaluations.
  • The 'critical' label in OpenAI's Preparedness Framework refers to models that can find or write functional zero‑day exploits or plan end‑to‑end attacks without human help, a step up from prior models rated at 'High.'
  • The disclosure heightens industry efforts to standardize testing and containment after recent containment failures from multiple labs and could prompt wider rules on controlled testing, reporting and defensive tools.