What 'Critical' Actually Means - and Why No Model Has Ever Crossed It
OpenAI's Preparedness Framework has carried a 'Critical' cybersecurity tier since its version-2 revision in April 2025, but until Astra it was a theoretical ceiling rather than a description of an actual model [5]. The bar is specific and severe: a Critical-rated system must be able to identify and develop functional zero-day exploits of all severity levels against hardened, real-world critical systems without human intervention, or independently devise and execute an entire novel cyberattack campaign against a hardened target from nothing more than a high-level goal [3]. That is a materially different capability from anything OpenAI has shipped. GPT-5.6 Sol, Terra and Luna - the company's most capable public models as of 2026 - all sit at the 'High' tier; GPT-5.6 Sol specifically could not produce a functional Critical-severity exploit against widely used hardened software under standard configuration [7]. Independent commentary outside OpenAI's own announcement backs that up: YouTube channel Smart AI Hustle noted that no prior OpenAI model, including GPT-5.6 Sol, had ever tripped the Critical cyber threshold before Astra, corroborating the 'first time' framing from outside the company's press materials.
What makes the Astra disclosure notable is not just the number on the scorecard, but that it is preliminary. OpenAI says benchmarking is still underway and it 'cannot rule out' the Critical designation - a hedge, not a confirmation - yet the framework's own rules required a precautionary response the moment that possibility became credible [2]. That is arguably the first real-world test of whether a voluntary internal safety commitment actually binds a lab's own roadmap, rather than existing only as a page on a website.


