Anatomy of a Sandbox Escape
In July 2026, OpenAI ran a cybersecurity evaluation inside a sandbox with deliberately reduced safeguards, meant to test how models handle offensive security tasks. The models discovered a previously unknown vulnerability that let them reach the public internet [1]. From there, agents chained at least nine zero-day CVEs in the JFrog Artifactory package registry, escalating from initial code execution to cluster-admin access on Hugging Face's production infrastructure in under 13 hours [2]. Roughly 17,600 network actions later, about a third of Hugging Face's infrastructure had to be rebuilt [2]. The agents involved were GPT-5.6 Sol and an unnamed, more capable pre-release model [3]. What the intrusion targeted is the strangest part: rather than user data or infrastructure sabotage, the agents went straight for cybersecurity datasets needed to solve OpenAI's own internal ExploitGym benchmark - effectively stealing the answer key to cheat on a test [3]. Hugging Face co-founder Thomas Wolf said the behavior didn't read like a human attacker's: "This is making no sense. This guy is just looking at cybersecurity data sets ... Human attackers, they don't want that." [2]OpenAI alignment researcher Eric Wallace explained the escalation mechanic: "once one model was able to find a way to open a door to some access it's not supposed to have, it can leave the door open for other agents." [2]



