Self-Reported Benchmarks vs Opus 4.8: A Split Decision, Not a Sweep
DeepSeek pitched V4-Flash-Vision-Exp, which went live on the DeepSeek API Platform on August 21, 2026, as a model that matches DeepSeek-V4-Flash on text capabilities while closing the multimodal-agent gap with Anthropic's Claude Opus 4.8 [1]. The headline numbers back that framing in places: on Agents' Last Exam, DeepSeek scored 27.3 against Opus 4.8's 25.7, and on ZeroBench it edged ahead 35.0 to 34.0 [5]. On ApexBench it landed close behind at 36.5 versus 39.4, and on Toolathlon-Verified and Chartography the two models finished within a point of each other [2]. Reporting from TheNextWeb puts the full picture in context, though: across the 11 benchmarks DeepSeek published, it actually wins only three, and the comparison itself was chosen carefully - DeepSeek benchmarked against Opus 4.8, a currently active model, rather than the newer Opus 5 [3].
The caveats run deeper than benchmark selection. The Decoder points out that DeepSeek's own evaluation was run through its internal Harness Minimal Mode, meaning the performance figures have not been independently verified [4]. Compounding that, when the underlying V4-Flash text model was forced to ignore the accompanying images, its ApexBench score collapsed to 26.2 - a reminder that the multimodal uplift is real, but the absolute numbers still come from a single, self-administered test harness rather than a neutral third party [5]. And the one gap DeepSeek doesn't close is arguably the one that matters most for paying customers: NL2Repo, a repository-scale coding benchmark, still shows Opus 4.8 ahead by 12 points (69.7 vs 57.7), which TheNextWeb calls a genuine competitive disadvantage for enterprise deployments [3].


