The Capability-Monitorability Tradeoff
GPT-6 Astra is the first OpenAI model to cross the 'Critical' cybersecurity capability threshold in the company's own Preparedness Framework [1][2]- meaning it can discover previously unknown security flaws and develop new exploits against well-protected systems without step-by-step human direction. That capability jump arrives paired with a change to how the model thinks: Astra uses a 'recurrent depth' (also called 'opaque recurrence') technique that lets it run extensive computation in loops rather than sequential, human-legible chain-of-thought steps [3], which reduces the reasoning traces available for outside inspection.
OpenAI's own Chief Scientist, Jakub Pachocki, has staked out a public commitment on where that tradeoff stops: "We will not accept degradation in our ability to monitor model alignment beyond a certain level. We will withhold scaling until we can regain enough confidence" [4]. Independent safety researchers are less reassured. Redwood Research CEO Buck Shlegeris warned that "If OpenAI pushes this technique further, they'll have the option to massively increase the recurrence and totally destroy CoT monitorability" [3], and former OpenAI safety lead Steven Adler went further, arguing that "OpenAI seems to be violating one of the few redlines that exist in the AI community" [5]. The disagreement is not about whether Astra is more capable - both sides agree it is - but about how much of its internal reasoning has become invisible, and how fast that opacity could scale.



