Three Wins Out of Eleven: The Benchmark Table Behind the Opus Comparison
DeepSeek's own launch announcement frames V4-Flash-Vision-Exp as closing the gap with Anthropic's Claude Opus 4.8 on multimodal agent tasks [1]. The model itself is not a from-scratch multimodal build - it is DeepSeek's existing V4-Flash backbone, a sparse mixture-of-experts architecture with 13B active parameters out of 284B total, extended with a vision encoder while text performance is held constant [1].
The fuller benchmark picture, once independent outlets tallied every number DeepSeek published, complicates the 'rivals Opus' framing. DeepSeek's model wins on Agents' Last Exam (27.3 vs 25.7), ZeroBench Pass@5 (35.0 vs 34.0) and DeepSWE (59.3 vs 58.0) - but loses on ApexBench Pass@1 (36.5 vs 39.4), Chartography (64.3 vs 65.0), and, most notably, NL2Repo, where the gap widens to 12 points (57.7 vs 69.7) [2]. OfficeChai's count puts it plainly: DeepSeek beats Opus 4.8 on 3 of the 11 tests it published [2]. That matters because the evaluation itself was run through DeepSeek's own internal 'Harness Minimal Mode' rather than an independently reproduced benchmark suite [2]- a caveat XenoSpectrum underscores directly, warning that small margins on a benchmark table are not proof the model can accurately read fine visual detail [3].



