The First-Ever "Critical" Verdict, and the Machinery Built to Contain It
OpenAI's internal evaluation of its next model, Astra, found strong enough agentic coding and cybersecurity performance that the company could not rule out reaching the "Critical" tier of its Preparedness Framework - the classification reserved for AI systems capable of independently finding or exploiting zero-day vulnerabilities and executing novel attacks against hardened systems [1]. That would make Astra the first model any AI lab has ever placed in that category [2]. OpenAI responded by pausing internal activities involving Astra that don't meet a new, stricter set of security requirements: isolated testing environments, restricted network and tool access, encrypted model weights, sandboxed execution, and real-time monitoring [3].
The monitoring isn't cosmetic. OpenAI has rolled out chain-of-thought surveillance across every agentic use of Astra, including during training and evaluation itself, layered with multistage automatic escalation that alerts safety, security, and research teams within 30 minutes of any flagged risky action, triggering an automatic pause if the alert isn't confirmed as a false alarm in that window [4]. That level of oversight isn't free: OpenAI estimates it adds roughly 20% additional compute overhead to Astra-related training and evaluation [4].



