Whose Blackwell? The Benchmark Fight Nobody's Won Yet
OpenAI's headline claim is stark: Jalapeño delivers 1.5x-1.9x more AI work per watt and 1.7x-3.6x lower latency than Nvidia's Blackwell-generation GB300 systems, running GPT-OSS 120B, DeepSeek's R1, and Moonshot AI's Kimi K2.5 [1]. On DeepSeek R1 specifically, OpenAI reported over 700 tokens per second per user at low concurrency, and close to 1,400 tokens per second per user on GPT-OSS - notably without using multi-token prediction, a common speed trick that rival systems reportedly did use in their own comparison numbers [2].
But the source of those numbers is also the source of the loudest caveat. SemiAnalysis, whose InferenceX benchmark suite generated the results, wrote in its own newsletter that comparing Jalapeño to Blackwell is 'somewhat incomplete and unfair,' since Nvidia's actual current-generation successor, Rubin, not Blackwell, would be the more applicable rival [2]. Community benchmarking discussion pushed the point further, noting that Jalapeño's own tests reportedly excluded speculative decoding - a technique that predicts several tokens ahead to save time - while OpenAI's Rubin comparisons apparently did include it, an asymmetry that, if real, would understate Jalapeño's true advantage once methodology is normalized. It also means nobody outside OpenAI has actually run that normalization. SemiAnalysis itself noted it hadn't independently executed the full InferenceX suite, nor seen results from AgentX, the benchmark meant to reflect real production agentic workloads [2]. Until someone outside OpenAI publishes numbers on the same models under the same conditions, 'beats Nvidia' is a claim resting entirely on the claimant's own test harness.


