OpenAI's Jalapeño Inference Chip
TECH

OpenAI's Jalapeño Inference Chip

36+
Signals

Strategic Overview

  • 01.
    OpenAI unveiled Jalapeno, its first custom AI inference chip, co-developed with Broadcom and fabricated by TSMC, presenting full technical details at Hot Chips 2026 on August 25, 2026.
  • 02.
    Benchmarks on SemiAnalysis's InferenceX suite, using OpenAI's GPT-OSS-120B plus DeepSeek R1 and Kimi K2.5, show 1.5x-1.9x higher throughput per kilowatt and 1.7x-3.6x lower end-to-end latency than Nvidia's GB200/GB300 systems.
  • 03.
    Each Jalapeno chip pairs a unified prefill-decode architecture with six HBM4 stacks (216 GiB, 15.4 TB/s bandwidth) inside a 700W package, versus 1,200W-1,400W for Nvidia and AMD's flagship GPUs.
  • 04.
    OpenAI plans to deploy Jalapeno in very small volumes in its own infrastructure by end of 2026, with broader rollout in 2027, while continuing to use Nvidia hardware alongside it.

Deep Analysis

A Clean-Sheet Chip Built to Move Data Less

A Clean-Sheet Chip Built to Move Data Less
OpenAI's Jalapeno inference chip, unveiled at Hot Chips 2026.

Jalapeno was designed from a blank sheet as an inference-only ASIC rather than a repurposed training chip. OpenAI has said plainly: "We designed Jalapeno to minimize data movement and communication delays." [1]The core idea is a unified prefill-decode architecture - instead of splitting the two phases of generating a response across separate hardware pools, as most GPU inference stacks do, Jalapeno handles both on the same silicon, with localized HBM sitting close to the compute that needs it. Each package carries six HBM4 stacks totaling 216 GiB of memory and 15.4 TB/s of bandwidth, and the design scales from 128-accelerator racks up to 2,048-chip pods delivering roughly 27 EFLOP/s of aggregate compute. [2]Getting there fast mattered too: OpenAI used its own models, including Codex, to help write and tune the chip's Gluon-language kernels - some hand-tuned to around 3,000 lines with a custom correctness sanitizer - which reportedly cut SIMD area by 8% and matrix-engine area by 10%, part of what let the team go from RTL freeze to tapeout in roughly nine months. [3]

The Real Currency Is Watts, Not FLOPs

The strongest argument for building custom inference silicon isn't raw compute - it's that inference, unlike training, is a recurring operating cost that scales with every user query, which makes power efficiency the thing that actually shows up on the bill. [4]Jalapeno's 700W package TDP (with sustained draw often measured lower during testing) compares against 1,200W for Nvidia's GB200 and 1,400W for GB300 and AMD's MI355X, which is where the reported 1.5x-1.9x throughput-per-kilowatt gains come from. Broadcom CEO Hock Tan has gone further, claiming Jalapeno runs inference at roughly half the cost of a typical AI GPU - a claim from Broadcom itself, not an independently audited figure. [5]Whether that cost advantage holds outside OpenAI's chosen test conditions is exactly the open question the rest of the industry is now arguing about.

A Benchmark OpenAI Graded Itself On

OpenAI ran its published Jalapeno numbers through SemiAnalysis's public InferenceX suite, testing its own GPT-OSS-120B alongside DeepSeek R1 and Moonshot AI's Kimi K2.5 against Nvidia's GB200 and GB300. [6]That is also the crux of the skepticism: SemiAnalysis's own writeup, despite concluding Jalapeno beats Blackwell in nearly every scenario tested, flags that OpenAI supplied all the benchmark data itself and that Blackwell may not even be the right comparison - Jalapeno's HBM4-based design more directly rivals Nvidia's newer Rubin chip, which wasn't tested at all. Community reaction on Reddit landed in roughly the same place: commenters broadly accepted the efficiency numbers as a genuine engineering achievement, but pushed back that this was effectively a vendor demo, run with OpenAI's engineers on-site and OpenAI's chosen models, on a workload (single-turn, fixed-length) that may not represent messier real-world agentic use. Several also noted the timing, landing the week ahead of Nvidia's earnings, as suspicious. On X, semiconductor analyst Dylan Patel - who said he got direct lab access to dissect the chip - called the result "huge news" and unusual for a first-generation chip, lending some independent credibility even as the underlying numbers remain OpenAI-run.

Chip War Optics, But OpenAI Is Still Buying Nvidia

The framing on social media has been unmistakably a rivalry story - OpenAI's own announcement post drew over 14,000 likes, and one widely-shared post claimed Nvidia's Jensen Huang had publicly responded to the comparison - a claim not independently confirmed elsewhere - which is itself a sign of how much this has been read as an OpenAI-versus-Nvidia contest rather than a narrow hardware upgrade. Analysts cited by CNBC said Jalapeno could pressure Nvidia's margins in the fast-growing inference segment and reduce OpenAI's reliance on Nvidia hardware for some workloads. [7]But the more understated fact is that OpenAI isn't dropping Nvidia at all: it plans to deploy Jalapeno in its own infrastructure only in very small volumes by the end of 2026, with more meaningful deployment in 2027, while continuing to buy Nvidia hardware in parallel. [8]That makes Jalapeno additive capacity and negotiating leverage first, wholesale replacement a much later and less certain story.

Shared Bottlenecks and a Multi-Generation Roadmap

Even if Jalapeno's numbers hold up under independent scrutiny, its near-term impact is capped by real supply constraints. TSMC's advanced packaging capacity - needed for Jalapeno's HBM4 stacks, reportedly sourced from Samsung - is sold out through 2026, a bottleneck OpenAI shares with Google and Meta's own custom-silicon efforts. [9]That mirrors Google's TPU program, which took roughly a decade of iteration before reaching a dedicated inference-specific chip in Ironwood - Jalapeno is arguably starting from a stronger position given its unusually fast nine-month design cycle, but it is still a first-generation part entering an industry where first-generation chips are rarely competitive out of the gate. OpenAI has said it is already working on Jalapeno's second and third generations, suggesting the real payoff, if there is one, will show up several chip cycles from now rather than in this initial small-volume rollout. [10]

Historical Context

~2016
Google began designing custom AI (TPU) chips roughly a decade before Jalapeno's unveiling, illustrating that custom-silicon programs typically take years of iteration; its seventh-generation TPU (Ironwood) is Google's first inference-specific chip.
2025-10
OpenAI and Broadcom signed a 10-gigawatt deployment agreement, the partnership announcement that preceded the Jalapeno unveiling.
2026-06-24
OpenAI and Broadcom first publicly unveiled the Jalapeno inference chip partnership and initial design goals.
2026-08-25
OpenAI's Richard Ho presented full Jalapeno architectural and benchmark details at the Hot Chips 2026 conference.

Power Map

Key Players
Subject

OpenAI's Jalapeño Inference Chip

OP

OpenAI

Chip designer and lead brand, sets the architecture and benchmarks, plans to deploy Jalapeno in its own data centers while remaining a customer of Nvidia hardware.

BR

Broadcom

Silicon co-design and implementation partner, also supplies the Tomahawk6 networking switches used in Jalapeno's rack fabric; CEO Hock Tan claims the chip runs inference at roughly half the cost of a typical AI GPU.

TS

TSMC

Fabricates the chip using advanced process and packaging; its advanced packaging capacity is reported sold out through 2026, a constraint shared with Google and Meta's own custom-silicon programs.

SA

Samsung Electronics

Reportedly supplies the HBM4 memory stacks used in the Jalapeno package.

NV

Nvidia

Incumbent GPU supplier and Jalapeno's primary benchmark rival via GB200/GB300; OpenAI continues to buy Nvidia hardware even as Jalapeno aims to reduce reliance on it for some inference workloads.

MI

Microsoft

Deployment partner - reported to receive Jalapeno deployment at gigawatt-scale data centers, alongside continued Azure infrastructure investment from OpenAI.

CE

Celestica

Handles board, rack, and system integration for Jalapeno deployments.

SE

SemiAnalysis

Independent analysis firm whose InferenceX benchmark suite was used for the published comparisons; concludes Jalapeno beats Blackwell broadly but notes the fairer long-term rival is Nvidia's upcoming HBM4-based Rubin chip.

Fact Check

10 cited
  1. [1] OpenAI's Jalapeno Chip Is Built for Fast Inference at Scale, Benchmarks Show
  2. [2] OpenAI Jalapeno ASIC at Hot Chips 2026
  3. [3] OpenAI Jalapeno: Better Than Nvidia Blackwell
  4. [4] Meet Jalapeno: OpenAI's First Custom AI Chip Built With Broadcom
  5. [5] Broadcom and OpenAI Unveil Custom-Built Jalapeno Inference Processor
  6. [6] First Benchmarks Revealed for Jalapeno, OpenAI's Clean-Sheet General Purpose AI Accelerator ASIC
  7. [7] OpenAI's Jalapeno AI Chip Raises Questions for Nvidia
  8. [8] OpenAI Details Jalapeno AI Chip With 700W TDP
  9. [9] OpenAI Debuts Jalapeno AI Inference Chip, With Samsung Reportedly Supplying HBM4
  10. [10] OpenAI Says Its Jalapeno Chip Beats Nvidia's GB300 in First Published Benchmarks

Source Articles

Top 5

THE SIGNAL.

Analysts

Presented Jalapeno's roughly nine-month RTL-to-tapeout timeline at Hot Chips 2026 and framed the results as a major jump over state-of-the-art hardware: "The bottom line is that the results show a very, very significant performance advance over state of the art."

Richard Ho
VP of Hardware, OpenAI

Argued the efficiency gain is real rather than marketing spin, saying it holds "on performance per watt and performance per dollar," and tied it to OpenAI's push to own more of its own infrastructure stack.

Greg Brockman
President and Co-Founder, OpenAI

Claims Jalapeno runs inference at "roughly half the cost of a typical AI GPU" - a claim from Broadcom itself, not an independently audited figure.

Hock Tan
CEO, Broadcom

Found Jalapeno "beats Blackwell... across almost all scenarios without being tuned for any specific point in the curve," excelling in both low-latency and high-throughput cases, while cautioning that OpenAI supplied all the benchmark numbers and that the fairer comparison is against Nvidia's HBM4-based Rubin chip, not Blackwell.

SemiAnalysis
Independent semiconductor analysis outlet
The Crowd

Since announcing Jalapeño, our first custom inference chip, we've been testing it and the system around it. The results show a major advance: more intelligence from every watt and faster responses, delivering both higher throughput and lower latency in one architecture without

@@OpenAI14162

OpenAI Jalapeno is spicy. Usually first generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Rubin. This is huge news! We got to go into OpenAI's lab to dissect their new chip and dive into the software, architecture, and performance. Incredible work!

@@dylan522p2130

WOW!!! Jensen of $NVDA responds to OpenAI claims that their jalapeño chip is better than Nvidia's GB300. SHOTS FIRED!!! Reminds me of You not talking to someone who woke up a loser moment!! What a boss!!

@@RealNickMugalli1049

OpenAI Just Dropped Benchmarks for Their Own Chip, Jalapeño, and It's Beating Nvidia's GB300

@u/AskGpts285
Broadcast
OpenAI Says New Jalapeno Chips Outperformed Nvidia in Testing

OpenAI Says New Jalapeno Chips Outperformed Nvidia in Testing

OpenAI's Jalapeño Chip Could Change LLMS Forever

OpenAI's Jalapeño Chip Could Change LLMS Forever

OpenAI's Jalapeño: The Chip That Ends the GPU Hype?

OpenAI's Jalapeño: The Chip That Ends the GPU Hype?