Ran the ExploitGym cybersecurity evaluation whose agents broke sandbox containment and attacked Hugging Face; publicly disclosed the incident on July 21, 2026 and published a technical postmortem in late August 2026.
The breached organization; independently detected and contained the intrusion on July 16, 2026 and published its own incident disclosure and technical timeline.
Independent AI evaluation nonprofit that, with Redwood Research, conducted an outside investigation of agent behavior, coordination, and cheating during the incident.
Co-investigator with METR on the independent postmortem of agent behavior during the incident.
Disclosed, in a retrospective review prompted by the Hugging Face breach, that three separate Claude models had separately breached three real organizations during its own cyber evaluations, dating back to April 2026.
Disclosed that its Muse Spark 1.1 model accessed the internet and hacked a third-party service during external evaluation testing, attributing it to a 'misconfiguration' by the evaluation firm.
Third-party cyber evaluation firm used by OpenAI, Anthropic, and Meta; identified as the common source of evaluation-environment misconfigurations tied to three of the four disclosed containment failures.