OpenAI's Jalapeño inference chip benchmarks vs Nvidia Blackwell
TECH

OpenAI's Jalapeño inference chip benchmarks vs Nvidia Blackwell

24+
Signals

Strategic Overview

  • 01.
    OpenAI presented benchmark results for Jalapeño, its custom inference ASIC built with Broadcom, at Hot Chips 2026 on August 25, 2026, claiming up to 1.9x higher performance-per-watt and up to 3.6x lower latency than Nvidia's GB200/GB300 Blackwell systems.
  • 02.
    Jalapeño is a 700W-TDP ASIC on TSMC's N3P node with 13.4 petaFLOPS of MXFP4 compute and 216 GB of HBM4 memory at 15.4 TB/s per chip, scaling to a 2,048-chip system delivering 27 exaFLOP/s and 432 TiB of memory.
  • 03.
    The chip is inference-only - OpenAI plans small-volume deployment in its own infrastructure by the end of 2026, with larger-scale deployment in 2027, while model training stays on Nvidia hardware.
  • 04.
    SemiAnalysis, whose InferenceX benchmark was used for the test, verified the runs in OpenAI's lab but relied on OpenAI-supplied numbers, and noted the Blackwell comparison is incomplete since Jalapeño's real rival generation is Nvidia's upcoming Rubin platform.

Deep Analysis

What Jalapeño Actually Claims - And What The Numbers Say

OpenAI published Jalapeño's first public benchmark results at Hot Chips 2026 on August 25, 2026, running the chip through SemiAnalysis's InferenceX suite rather than a benchmark of its own design [1]. The chip itself is a 700-watt-TDP ASIC fabricated on TSMC's N3P process, packing 13.4 petaFLOPS of MXFP4 matrix compute per chip with 216 GB of HBM4 memory delivering 15.4 TB/s of bandwidth; a single rack holds 128 chips, and OpenAI's full reference system scales to 2,048 chips for 27 exaFLOP/s and 432 TiB of pooled memory [2]. On the two flagship workloads OpenAI published, GPT-OSS 120B and DeepSeek R1 670B, Jalapeño posted 1.9x and 1.7x higher peak mixed tokens-per-second-per-kilowatt than Nvidia's GB200, alongside 1.7x and 3.6x lower end-to-end latency [2].

The efficiency gap traces to a specific architectural choice: OpenAI designed the chip to minimize data movement so that model state, including the KV cache generated while a response streams out, can be explicitly placed and kept local rather than shuttled across the system [3]. That is what let OpenAI claim something GPUs traditionally struggle to deliver at once - high throughput and low latency on the same hardware - a result OpenAI's head of hardware, Richard Ho, described as a very significant performance advance over the state of the art [1]. A later B0 silicon stepping is reported to add roughly another 25% performance-per-watt over the earlier A0 silicon [4].

A Compressed Design Cycle - the Meta-Story Behind the Silicon

Jalapeño's own disclosed timeline is almost as notable as its benchmark scores: architecture concept work began in late 2024, an RTL freeze followed in 2025, tapeout landed in late 2025, and OpenAI's Codex model was reportedly running on early silicon by early 2026 [2]. That is an unusually compressed cycle for a from-scratch inference ASIC, and it fits the broader positioning OpenAI staked out when it first unveiled the Broadcom partnership: that OpenAI is no longer just a model company but is building the infrastructure underneath its products, down to chip architecture, kernels, memory systems, networking, and scheduling [5].

Social commentary around the Hot Chips talk framed the speed itself as part of the story, with engineers describing an AI-assisted co-design loop and AI-generated kernel code that reportedly beat hand-tuned implementations on some critical blocks. Whether or not that fully explains the compressed schedule, it lines up with OpenAI's stated rationale for building Jalapeño at all: designing for one company's own well-understood, repeatable inference workload lets a much smaller team move faster than a general-purpose GPU roadmap built to serve every customer at once.

The Asterisks: Vendor Numbers, a Missing Rival, and an Uneven Fight

The headline beats-Nvidia framing comes with real caveats that SemiAnalysis - the firm whose own InferenceX benchmark was used - was careful to flag. All of the published numbers came from OpenAI; SemiAnalysis says it verified the InferenceX runs in person in OpenAI's lab but did not independently execute the full benchmark suite itself [6]. More pointedly, SemiAnalysis noted the comparison to Blackwell is somewhat incomplete and unfair, since Jalapeño's real competitor generation is Nvidia's upcoming Rubin platform [6]. SemiAnalysis's own read, however, was still that Jalapeño beats Blackwell on performance-per-watt across almost all scenarios without being tuned for any specific point [6].

That same tension - real efficiency gains, measured against a benchmark that skips the actual next-generation rival - showed up in community reaction, where the chip was treated like a vendor demo until someone independent reproduces it, with recurring skepticism that excluding Rubin stacks the comparison in OpenAI's favor. A separate thread questioned whether an ASIC tuned for today's transformer-based inference risks aging out as model architectures shift, a concern OpenAI's own framing implicitly answers by targeting general transformer-style inference as a stable workload rather than betting on one model family.

Why Now: Nvidia Dependence, Margins, and Timing

The strategic logic behind Jalapeño predates the benchmarks by two months: OpenAI and Broadcom first revealed the partnership on June 24, 2026, targeting roughly 10 gigawatts of custom chip capacity for gigawatt-scale data center deployment [7]. The underlying motive is cost and leverage - chipmakers like Nvidia have historically commanded profit margins as high as 75% on AI hardware, and even a modest volume of self-designed inference chips strengthens OpenAI's negotiating position with its suppliers [8]. Inference is also simply a better fit for custom silicon than training: the workload is large, stable, and repeatable at OpenAI's scale, which is why Jalapeño is explicitly not designed to train models at all [9].

Analysts have largely read the results as significant for Nvidia's economics; one Yole Group analyst argued the results show a hyperscaler-designed chip can now match or beat Nvidia's Blackwell-class GPUs on inference efficiency, calling it a real threat to Nvidia's fast-growing inference margins even though Nvidia retains most AI compute share and CUDA lock-in [8]. Nvidia CEO Jensen Huang publicly waved off the threat, noting that lots of projects get started and lots of projects get canceled, while pointing to Nvidia's own accelerating market share and revenue as proof its position isn't at risk [10]. For now, the practical impact is limited: initial Jalapeño production is planned in small volumes by the end of 2026, with the bulk of OpenAI's inference traffic staying on Nvidia hardware for roughly the next 18 months [8].

Historical Context

2024
Jalapeño's architecture concept phase began in late 2024, ahead of a 2025 RTL freeze and late-2025 tapeout.
2026-06-24
OpenAI and Broadcom publicly unveiled Jalapeño, OpenAI's first custom AI chip, targeting roughly 10 gigawatts of custom chip capacity.
2026-08-25
OpenAI published Jalapeño's first benchmark results at Hot Chips 2026, claiming 1.5x-1.9x higher performance-per-watt and lower latency versus Nvidia GB200/GB300 Blackwell systems.

Power Map

Key Players
Subject

OpenAI's Jalapeño inference chip benchmarks vs Nvidia Blackwell

OP

OpenAI

Designed Jalapeño with Broadcom as its first custom AI silicon to cut Nvidia dependence for inference and gain supplier leverage, with small-volume deployment by end of 2026 and a larger rollout in 2027.

BR

Broadcom

Co-designs and manufactures Jalapeño with OpenAI, framing it as the first generation of a multi-year compute platform targeting roughly 10 gigawatts of custom chip capacity.

NV

Nvidia

Incumbent GPU supplier whose Blackwell (GB200/GB300) systems are the benchmark baseline Jalapeño claims to beat; CEO Jensen Huang publicly downplayed the threat while citing accelerating market share and revenue.

SE

SemiAnalysis

Independent semiconductor analysis firm whose InferenceX benchmark was used to test Jalapeño; verified runs in-lab but flagged the comparison as favorable to OpenAI and incomplete versus Nvidia's next-gen Rubin.

TS

TSMC

Manufactures the Jalapeño compute die on its N3P process node.

Fact Check

10 cited
  1. [1] OpenAI's Jalapeño Chip Is Built For Fast Inference At Scale, Benchmarks Show
  2. [2] OpenAI Jalapeno ASIC At Hot Chips 2026
  3. [3] OpenAI's Upcoming Jalapeno Chip Looks Like It'll Be An Inference Beast
  4. [4] OpenAI's Jalapeno Chip Outperforms Nvidia's Blackwell In Efficiency
  5. [5] OpenAI Unveils Its First Custom Chip Built By Broadcom
  6. [6] OpenAI Jalapeño: Better Than Nvidia Blackwell
  7. [7] OpenAI And Broadcom Reveal Jalapeno, First AI Chip In Partnership
  8. [8] OpenAI's Jalapeno AI Chip Takes Aim At Nvidia
  9. [9] How OpenAI's Jalapeno Chip Changes Nvidia's Economic Picture
  10. [10] Nvidia's Jensen Huang On OpenAI's Jalapeno Chip Competition

Source Articles

Top 5

THE SIGNAL.

Analysts

Says the Jalapeño results show a hyperscaler-designed chip can now match or beat Nvidia's Blackwell-class GPUs on inference efficiency, calling it a real threat to Nvidia's fast-growing inference margins even though Nvidia keeps most AI compute share and CUDA lock-in.

Adrien Sanchez
Technology analyst, Yole Group

Downplayed Jalapeño as unremarkable in an industry where many custom chip projects start and get cancelled, pointing to Nvidia's growing market share and accelerating revenue as evidence its position isn't threatened.

Jensen Huang
CEO, Nvidia

Confirms Jalapeño beats Blackwell on performance-per-watt across almost all tested scenarios without being tuned for any specific point, but cautions the comparison is somewhat unfair since Jalapeño's real rival generation is Nvidia's upcoming Rubin, and OpenAI supplied the underlying numbers.

SemiAnalysis
Independent semiconductor analysis firm (InferenceX benchmark)
The Crowd

Since announcing Jalapeño, our first custom inference chip, we've been testing it and the system around it. The results show a major advance: more intelligence from every watt and faster responses, delivering both higher throughput and lower latency in one architecture without sacrificing efficiency.

@@OpenAI14058

OpenAI says its first custom chip just beat Nvidia's GB300. Jalapeño: up to 1.9× more AI work/watt. Up to 3.6× lower latency. Built in 9 months. Then Nvidia reported a $96B quarter. The chip war just became a benchmark war.

@@JonhernandezIA19

All the focus around Jalapeño seems centered on performance vis-a-vis Nvidia, but the engineering velocity behind it is truly remarkable. Some highlights from the @hotchipsorg talk: 1. Building custom chips used to require massive armies of engineers and multi-year timelines. OpenAI went from initial RTL code to tapeout in just 9 months using XLS (a Rust-like language for hw from Google) and AI-driven co-design. 2. An AI loop took a raw, barely functional (DeepSeek) kernel running at 0.31% efficiency and autonomously optimized it to 88.94% of the chip's physical limit in under 2 days. 3. The AI-generated code ran up to 1.8x faster on critical blocks than implementations hand-tuned by world-class kernel engineers. 4. AI did more than just write software for the chip. It optimized the physical hardware circuits before tapeout.

@@ai51

OpenAI Just Dropped Benchmarks for Their Own Chip, Jalapeño, and It's Beating Nvidia's GB300

@u/AskGpts283
Broadcast
OpenAI Says New Jalapeno Chips Outperformed Nvidia in Testing

OpenAI Says New Jalapeno Chips Outperformed Nvidia in Testing

OpenAI launched Jalapeno, their own AI chip!

OpenAI launched Jalapeno, their own AI chip!

OpenAI's Jalapeño Chip Could Change LLMS Forever

OpenAI's Jalapeño Chip Could Change LLMS Forever

OpenAI's Jalapeño inference chip benchmarks vs Nvidia Blackwell — AI News | Agentic Brew