OpenAI's Jalapeño inference chip challenges Nvidia
TECH

OpenAI's Jalapeño inference chip challenges Nvidia

40+
Signals

Strategic Overview

  • 01.
    OpenAI and Broadcom unveiled Jalapeño, OpenAI's first custom AI chip, an 'Intelligence Processor' built specifically for LLM inference rather than training.
  • 02.
    The chip moved from early hiring and schematics to tape-out readiness in roughly nine months by one account, or about sixteen months by another - unusually fast for a semiconductor program.
  • 03.
    OpenAI claims Jalapeño delivers 1.5x-1.9x better performance per watt and 1.7x-3.6x lower end-to-end latency than Nvidia Blackwell/GB300 systems on GPT-OSS 120B, DeepSeek R1, and Kimi K2.5, based on SemiAnalysis' InferenceX benchmark suite.
  • 04.
    Jalapeño uses HBM4 memory with 15.4TB/s of bandwidth and an architecture that avoids splitting prefill and decode across separate hardware pools.
  • 05.
    Broadcom contributed silicon implementation and Tomahawk networking technology; Celestica assisted with board, rack, and system integration for deployment.
  • 06.
    Deployment plan: a small-volume prototype rollout by end of 2026, with more substantial volume production and deployment in 2027.
  • 07.
    The chip is rated at 700 watts during inference testing and was benchmarked against Nvidia's GB300, not Nvidia's newer Vera Rubin generation.
  • 08.
    A second-generation Jalapeño chip is already in advanced development, with tape-out expected soon.

Deep Analysis

Whose Blackwell? The Benchmark Fight Nobody's Won Yet

OpenAI's headline claim is stark: Jalapeño delivers 1.5x-1.9x more AI work per watt and 1.7x-3.6x lower latency than Nvidia's Blackwell-generation GB300 systems, running GPT-OSS 120B, DeepSeek's R1, and Moonshot AI's Kimi K2.5 [1]. On DeepSeek R1 specifically, OpenAI reported over 700 tokens per second per user at low concurrency, and close to 1,400 tokens per second per user on GPT-OSS - notably without using multi-token prediction, a common speed trick that rival systems reportedly did use in their own comparison numbers [2].

But the source of those numbers is also the source of the loudest caveat. SemiAnalysis, whose InferenceX benchmark suite generated the results, wrote in its own newsletter that comparing Jalapeño to Blackwell is 'somewhat incomplete and unfair,' since Nvidia's actual current-generation successor, Rubin, not Blackwell, would be the more applicable rival [2]. Community benchmarking discussion pushed the point further, noting that Jalapeño's own tests reportedly excluded speculative decoding - a technique that predicts several tokens ahead to save time - while OpenAI's Rubin comparisons apparently did include it, an asymmetry that, if real, would understate Jalapeño's true advantage once methodology is normalized. It also means nobody outside OpenAI has actually run that normalization. SemiAnalysis itself noted it hadn't independently executed the full InferenceX suite, nor seen results from AgentX, the benchmark meant to reflect real production agentic workloads [2]. Until someone outside OpenAI publishes numbers on the same models under the same conditions, 'beats Nvidia' is a claim resting entirely on the claimant's own test harness.

Nine Months, Not Nine Years

Custom AI silicon typically takes years to go from concept to working hardware. Jalapeño reportedly moved from early hiring and schematics to a fabrication-ready, reticle-sized ASIC in roughly nine months [3], or about sixteen months counting from initial hiring to tape-out by a more conservative account [2]. Either way, it is an unusually compressed cycle for the industry.

The speed came from an aggressive form of software-hardware co-design that used OpenAI's own models during development, including generating kernels through its coding tools [4]. Architecturally, the chip skips a common industry pattern - splitting prefill (processing the incoming prompt) and decode (generating the response token by token) across separate hardware pools - and instead keeps both on unified resources [2], specifically to cut the data-movement and communication delays that usually bottleneck inference at scale [1]. That design choice is a bet that avoiding hardware specialization overhead matters more than optimizing each phase separately, and it's a bet OpenAI could only make because it controls both the model workloads and the chip design simultaneously - something a merchant silicon vendor selling to many customers with different workloads cannot easily do.

Not an Nvidia Killer - A Margin Defense System

The plainest explanation for why Jalapeño exists came from OpenAI's own hardware lead: 'We have so much need for compute' [5]. That shortage, not a strategic desire to displace Nvidia outright, is the stated driver - OpenAI's compute demand has outstripped what the general-purpose GPU market can supply on its own terms. Greg Brockman framed the payoff in economic terms, calling it 'a real performance improvement...on performance per watt and performance per dollar' [4], and multiple outlets reported the efficiency gains translate to roughly a 50 percent cut in inference costs [4].

But the scope is narrower than the framing sometimes suggests. Jalapeño only handles inference, not training, and OpenAI's overall compute stack stays reliant on Nvidia GPUs (and other suppliers) for training and a large share of inference workloads well into 2027 [1]. Coverage has described the chip as putting 'Nvidia's pricing power on notice' [6], which is a more accurate read than 'Nvidia killer': Jalapeño gives OpenAI a credible alternative to point to in negotiations and a hedge against GPU supply constraints, without actually replacing the bulk of what it buys from Nvidia. It's optionality, not independence.

The 2027 Execution Gap

Even by OpenAI's own account, Jalapeño's near-term footprint is small. The company describes an initial rollout at very small volumes by the end of 2026, with more substantial deployment not expected until 2027 [1], and a second-generation chip is already in advanced development with tape-out expected soon [5]- suggesting OpenAI is iterating faster than it is shipping at scale.

Skepticism online has centered on whether that 2027 volume target is realistic at all, given the ongoing HBM memory shortage and Nvidia's entrenched grip on advanced packaging and fab capacity that any competing chip program also has to draw on. That skepticism sits alongside a broader, more structural question raised in the same discussions: whether an ASIC with better raw performance-per-watt can meaningfully dent Nvidia's position when Nvidia's actual moat is arguably its software and networking ecosystem - CUDA, interconnects, tooling - rather than silicon efficiency alone. Jalapeño's benchmark numbers, however they shake out under independent scrutiny, don't resolve that question; they just raise the stakes for whether OpenAI can actually manufacture and deploy at the volume it's promising.

Historical Context

2025-10
OpenAI and Broadcom officially announced their chip development partnership.
2026-06-24
OpenAI and Broadcom publicly unveiled the Jalapeño chip for the first time as OpenAI's first custom AI inference chip.
2026-08-25
OpenAI released first benchmark results for Jalapeño via SemiAnalysis' InferenceX suite, claiming 1.5x-1.9x better performance per watt and 1.7x-3.6x lower latency versus Nvidia Blackwell/GB300.

Power Map

Key Players
Subject

OpenAI's Jalapeño inference chip challenges Nvidia

OP

OpenAI

Designs the chip architecture and drives a full-stack compute strategy to reduce dependence on Nvidia for inference specifically, while still relying on Nvidia and others for training - if OpenAI hadn't pursued this, it would remain a pure GPU-market price-taker for its highest-volume workload.

BR

Broadcom

Co-develops and manufactures the chip, supplying silicon implementation and Tomahawk networking - Broadcom's fabrication capacity is what turns OpenAI's chip designs into physical, deployable hardware.

NV

Nvidia

Remains the benchmark comparison point (GB300/Blackwell) and OpenAI's supplier for training and other inference work; Jalapeño only threatens Nvidia's pricing leverage on a slice of inference demand, not its core training franchise.

CE

Celestica

Handles board, rack, and system integration for Jalapeño deployment - a less visible but necessary link between chip design and actual data-center rollout.

SE

SemiAnalysis

Independent analyst firm whose InferenceX benchmark suite produced the headline numbers, while its own newsletter simultaneously questioned the fairness of the comparison - making it both the source of OpenAI's proof and its most credible skeptic.

Fact Check

6 cited
  1. [1] OpenAI's Jalapeño chip is built for fast inference at scale, benchmarks show
  2. [2] OpenAI Jalapeño: Better Than Nvidia?
  3. [3] Broadcom and OpenAI Unveil Custom-Built Jalapeño Inference Processor
  4. [4] OpenAI Unveils First Custom AI Inference Chip Jalapeño With Broadcom, and Its Development Was Sped Up With OpenAI's Own Models
  5. [5] OpenAI Jalapeño vs Nvidia GB300 Benchmarks
  6. [6] OpenAI and Broadcom Unveil OpenAI Jalapeño, a Custom Inference Chip That Puts Nvidia's Pricing Power on Notice

Source Articles

Top 5

THE SIGNAL.

Analysts

Frames the benchmark results as a major, credible performance advance that delivers both higher throughput and lower latency at once, tying the effort back to OpenAI's compute demand outstripping available supply.

Richard Ho
Head of OpenAI's chip/hardware division

Frames the results in concrete economic terms, calling it a genuine performance improvement rather than a marginal one, specifically on performance per watt and performance per dollar.

Greg Brockman
President and co-founder, OpenAI

Cautions that comparing Jalapeño to Blackwell rather than Nvidia's upcoming Rubin is incomplete.

SemiAnalysis (newsletter analysis)
Independent semiconductor/AI infrastructure analysis outlet
The Crowd

Since announcing Jalapeño, our first custom inference chip, we've been testing it and the system around it. The results show a major advance: more intelligence from every watt and faster responses, delivering both higher throughput and lower latency in one architecture without sacrificing efficiency.

@@OpenAI3187

OpenAI Jalapeno is spicy Usually first generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Rubin This is huge news! We got to go into OpenAI's lab to dissect their new chip and dive into the software, architecture, and performance Incredible work!

@@dylan522p1095

Holy: OpenAI says its first custom inference chip is already beating Nvidia GB200 and GB300 systems on speed and efficiency. Across GPT‑OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, Jalapeño delivered 1.5–1.9× more AI work per watt at peak throughput and 1.7–3.6× lower end-to-end latency in OpenAI's own InferenceX testing! For highly interactive workloads, OpenAI reports 2.1–4.1× higher performance. The chip is rated at 700 watts, but remained at or below 550 watts during the tested workloads. OpenAI plans to begin deploying Jalapeño by the end of 2026. Gen 2 is already deep in development, with Gen 3 taking shape.

@@kimmonismus1000

OpenAI's new chip is better than Vera rubin on benchmark

@u/Wonderful_Buffalo_32102
Broadcast
OpenAI’s Jalapeño Chip Could Change LLMS Forever

OpenAI’s Jalapeño Chip Could Change LLMS Forever

OpenAI launched Jalapeno, their own AI chip!

OpenAI launched Jalapeno, their own AI chip!

OpenAI’s first AI chip is called Jalapeño #Vergecast

OpenAI’s first AI chip is called Jalapeño #Vergecast

OpenAI's Jalapeño inference chip challenges Nvidia — AI News | Agentic Brew