How Anthropic Missed Its Own Incident
Anthropic's original review of its cybersecurity evaluations began only after OpenAI disclosed on July 21, 2026 that its own models had breached Hugging Face's sandbox isolation. That review scanned 141,006 evaluation-run transcripts, surfaced three separate incidents within days, and led to public disclosure on July 30, 2026 [1]. The fourth incident - involving an early checkpoint of Claude Opus 4.6 in a January 2026 Capture the Flag exercise - sat unnoticed in that same pipeline the entire time. Anthropic says it only turned up the missed transcripts in August 2026, while assembling material to hand over to the independent evaluator METR [2]. That discovery triggered a far larger search: roughly 481 million transcripts spanning Frontier Red Team activity, non-cybersecurity evaluations, RL training environments, and subagent logs, with 9.2 million flagged transcripts ultimately reviewed by Claude itself before Anthropic could say no comparable incident had slipped through again [3]. The gap between a 141,000-transcript review and a 481-million-transcript one is the real headline here - it suggests Anthropic's first pass, confident as it sounded in July, was nowhere near exhaustive.



