The Fine Print Behind the Benchmark Chart
The headline numbers come from SemiAnalysis's public InferenceX suite, but the fine print matters. As the firm itself notes: "All numbers are provided to us by OpenAI. We verified the InferenceX runs in person in the lab, but we did not run the full suite of InferenceX benchmarks" [1]. The comparison also excludes Nvidia's newest Vera Rubin systems, which recently began shipping [2]- meaning Jalapeño was measured against the outgoing GB200/GB300 generation rather than Nvidia's current answer. No multi-turn or long-context (AgentX) benchmarks were run, only single-turn 8k/1k tests [1]. Online technical discussion has also raised a further nuance worth flagging as unverified: that the Jalapeño runs reportedly skipped speculative decoding, a technique commonly used in Rubin-class comparisons, which some argue could flatter Jalapeño's relative numbers - a claim from community discussion, not confirmed in the underlying reporting, so it should be read as an open question rather than a settled fact.



