The 95% Completion-Rate Headline Measures Compliance, Not Skill
The benchmark numbers are dramatic on their face. GPT-5.5 posted a 71.4% (±8.0%) pass rate on the UK AI Security Institute's expert-level cyber tasks, ahead of Claude Mythos Preview's 68.6% and OpenAI's own prior release, GPT-5.4, at 52.4%[1]. On Irregular's atomic challenge suite, the same model hit 98% on network attack simulation challenges and 92% on vulnerability research and exploitation challenges[2]. OpenAI's newer, more specialized GPT-5.6-Cyber then reported 95% completion on the company's own Advanced Cybersecurity Completion Rate benchmark, versus 57.3% for GPT-5.5-Cyber and just 1.5% for the standard GPT-5.6 Sol model with its normal safety guardrails applied[3].
That GPT-5.5-versus-Mythos comparison became a talking point outside the benchmark write-ups too. On r/ArtificialInteligence, a user who said they'd used GPT-5.5 hands-on to hunt for vulnerabilities weighed in directly on Anthropic's own "too dangerous to release" language about Mythos, verified via the same AI Security Institute running these cyber evaluations. Their verdict after actually using the model for vulnerability research: it's "pretty good... but hardly 'too dangerous to release.'"
That last comparison is the tell. VentureBeat's analysis of the release put it bluntly - the completion-rate metric measures how often a model attempts a task, not whether it gets the task right, calling it 'a refusal metric wearing a capability metric's clothes'[4]. The jump to 95% looks less like a leap in raw skill and more like OpenAI dialing back the guardrails that held the base model at 1.5%. Real capability gains are visible elsewhere, though: GPT-5.6-Cyber reportedly surfaced two previously unknown vulnerabilities in Chrome's V8 engine (one assigned CVE-2026-15903), several bugs in a mobile OS, and more than 400 privilege-escalation bugs in a popular OS kernel[3].
Open-weight models are closing the gap from a different angle. DeepSeek-V4-Flash scored 76.7 on the Cybergym security benchmark, roughly double the 38.7 its preview version managed[5], and DeepSeek shipped a distilled variant, DeepSeek-V4-Fable, purpose-built for autonomous CTF-style security research and multi-step exploitation planning[6]. Unlike OpenAI's gated Daybreak Red tier, these models are openly downloadable - the capability curve isn't just rising, it's diffusing.


