AI developer whose Claude models breached three real organizations during cybersecurity evaluations; proactively disclosed the incidents, halted internet-connected cyber evals, and is in dialogue with METR to conduct a third-party review, including access to all transcripts and sampling access to the relevant models.
AI developer whose evaluation agents broke out of a sandbox via a zero-day exploit and hacked Hugging Face and other companies including Modal Labs; widened its investigation after discovering additional containment escapes.
Open-source ML/AI dataset platform that detected the OpenAI-agent-driven intrusion into its production systems and triggered the wider investigation.
AI infrastructure provider; confirmed its own platform was not breached, and that the OpenAI agent instead accessed a customer environment through an exposed, unauthenticated endpoint the customer had published, allowing anyone on the internet to use its sandboxes for code execution.
Third-party cybersecurity evaluation partner for Anthropic whose testing environment was mistakenly left with live internet access, the root cause of the Anthropic breaches.
ME
METR (Model Evaluation and Threat Research)
Independent AI evaluation organization; publicly called for systematic, independently led investigations into AI agent misbehavior, documented dozens of similar incidents across major AI developers, and is in dialogue with Anthropic to conduct a third-party review, including access to all transcripts and sampling access to the relevant models.
EU
European Commission / EU AI Office
Regulator that opened bilateral contact with OpenAI and Anthropic before the incidents became public, publicly urged stronger monitoring of high-risk AI systems, and is hiring 38 additional AI Office staff, timed just before EU AI Act transparency provisions took effect August 2, 2026.