OpenAI's Jalapeño custom inference chip debut
TECH

OpenAI's Jalapeño custom inference chip debut

32+
Signals

Strategic Overview

  • 01.
    OpenAI and Broadcom unveiled Jalapeño, OpenAI's first custom AI accelerator - a purpose-built 'Intelligence Processor' designed for LLM inference.
  • 02.
    First published benchmarks show Jalapeño delivering 1.5x-1.9x higher peak performance-per-watt and 1.7x-3.6x lower end-to-end latency than Nvidia's GB200/GB300 systems across GPT-OSS-120B, DeepSeek R1, and Kimi K2.5.
  • 03.
    A single Jalapeño chip packs 13.4 PFLOPS of MXFP4 compute, 216 GiB of HBM4 memory, and 15.4 TB/s of bandwidth at 700W, scaling to 27 EFLOPS and 432 TiB of memory across a 2,048-chip pod.
  • 04.
    Jalapeño is inference-only (no training support) and will deploy in OpenAI's own infrastructure in very small volumes by the end of 2026, with more substantial production expected in 2027.

Deep Analysis

The AI Designed Its Own Chip

OpenAI's chip effort moved from initial design to tapeout in an estimated nine to sixteen months [5], an unusually fast timeline that one industry analysis flagged as a possible record for chip design [1]- compressed largely because OpenAI turned its own models loose on the hardware itself. AI-written kernel implementations for Jalapeño outperformed human expert kernels by 1.5 to 1.8 times [2], with AI also used to accelerate verification loops and low-level kernel optimization throughout the design process [2]. A technical teardown found the AI-tuned matrix units delivered roughly a 56% gain on BF16 multiply operations while shrinking matrix-unit die area by about 10% [3]. In a Bloomberg Technology interview, Richard Ho pointed to a different, faster-moving marker of that speed: three very different AI models were brought up and running on the new hardware within a couple of months of OpenAI receiving it, which he cited as evidence that the programming model itself had matured fast enough to keep pace with the silicon. The resulting single chip packs 13.4 PFLOPS of MXFP4 compute, 216 GiB of HBM4 memory, and 15.4 TB/s of bandwidth in a 700W envelope [3], scaling to 27 EFLOPS and 432 TiB of memory across a 2,048-chip pod [3]. The notable part isn't just that a lab built a chip - it's that the chip's own creation leaned on the AI it exists to run, folding model-assisted engineering into a discipline that historically took human teams years.

Why the Benchmark Comparison Is Already Contested

OpenAI's headline numbers - 1.5x-1.9x higher throughput per watt and 1.7x-3.6x lower end-to-end latency than Nvidia's GB200/GB300 systems [4]- sound decisive, but the fine print complicates the story. SemiAnalysis, which independently reviewed the InferenceX comparisons, called the results 'industry-leading' for a first-generation chip [5], while also cautioning that measuring Jalapeño against Blackwell is somewhat incomplete and unfair, since Nvidia's upcoming HBM4-based Vera Rubin platform is the more contemporaneous rival, and OpenAI's tests used single-token prediction where Rubin comparisons typically use speculative decoding [5]. Richard Ho, OpenAI's head of hardware, framed the results as 'a very, very significant performance advance over state of the art' [6]. Separately, in a Bloomberg Technology interview, Ho noted the benchmarks were run on InferenceX, an open-source benchmarking tool chosen specifically so the comparison wouldn't rest on OpenAI's own scoring. Even so, no independent full technical report has been published yet, and that gap between OpenAI's framing and the benchmark's actual scope is exactly where the loudest skepticism lives: the comparison is against a two-year-old architecture, not the chip Nvidia is about to ship. Discussion across r/OpenAI and r/tomshardware read as impressed but cautious, and the same Rubin-skepticism surfaced independently there too - a widely upvoted r/tomshardware thread (224 upvotes) made a point of noting Vera Rubin's absence from the comparison, arriving at the same conclusion as SemiAnalysis without citing it. A separate, heavily upvoted Reddit comment pushed the critique further, arguing Jalapeño's real edge is concentrated in low-latency scenarios rather than raw total throughput - a distinction OpenAI's headline multiples tend to blur.

The Economics of Inference Independence

Broadcom CEO Hock Tan pegs the accelerator's cost advantage at roughly 50% savings versus typical AI GPUs [7], and OpenAI designed Jalapeño specifically to minimize data movement and communication delays across the prefill and decode phases of LLM serving [6]- the parts of inference that dominate cost at OpenAI's scale. But the economics only matter once the chip ships in volume, and that is still a while off: deployment in OpenAI's own infrastructure starts in very small volumes by the end of 2026, with more substantial production not expected until 2027 [6]. That timeline means Jalapeño's near-term effect on OpenAI's Nvidia spend, and on Nvidia's roughly 92% share of the GPU market [8], is limited - the announcement reads as a statement of intent and a proof of technical capability more than an immediate supply-chain shift. Reddit commentary converged on the same point from a different angle: several commenters noted Jalapeño is still 'in early production' while Nvidia's Rubin cards are already being produced in the hundreds of thousands to millions, a volume gap that will matter more to near-term economics than any benchmark multiple.

Jensen Huang's Calculated Shrug

Nvidia CEO Jensen Huang's public response leaned on dismissal rather than direct rebuttal: 'Lots of projects get started. Lots of projects get canceled,' he said, framing custom-silicon efforts like Jalapeño as a common and often unsuccessful pattern among Nvidia's largest customers [8]. Reporting also framed the chip as a new pressure point on Nvidia's margins specifically in the inference segment, where purpose-built ASICs are best positioned to compete [9]. Huang's response is notable less for what it disputes and more for what it doesn't: it doesn't contest OpenAI's benchmark numbers, it bets instead that Nvidia's platform breadth, and the sheer difficulty of taking a first-generation chip to reliable volume production, will outlast one strong showing at Hot Chips. On X, OpenAI's own two announcement and benchmark posts drove far more engagement than any independent commentary - 22,544 and 14,127 likes respectively - underscoring that this is a story the company is actively driving rather than one that emerged organically from outside reaction. Independent voices did weigh in: a finance-focused account, @RealNickMugalli, reacted to Huang's remarks with 'SHOTS FIRED,' framing the response as combative bravado rather than a substantive rebuttal - a read that lines up with how little Huang's comments actually engaged the benchmark numbers themselves.

Historical Context

2025-10-13
OpenAI and Broadcom announced a strategic collaboration to deploy 10 gigawatts of OpenAI-designed AI accelerators.
2025-11
Jalapeño's design taped out for manufacturing, an estimated nine to sixteen months after initial development work began.
2026-06-24
OpenAI and Broadcom publicly unveiled the Jalapeño chip design and announced OpenAI had received first silicon samples.
2026-08-25
OpenAI published Jalapeño's first benchmark results around Hot Chips 2026, claiming 1.5x-1.9x throughput-per-watt and 1.7x-3.6x latency advantages over Nvidia GB200/GB300.
2026-08-26
Nvidia CEO responded to Jalapeño publicly, downplaying the chip as a competitive threat.

Power Map

Key Players
Subject

OpenAI's Jalapeño custom inference chip debut

OP

OpenAI

Designer of Jalapeño and its target deployer, seeking to reduce reliance on Nvidia GPUs and control its own inference cost and latency economics.

BR

Broadcom

Co-development partner handling silicon implementation and networking; CEO Hock Tan claims roughly 50% cost savings versus typical AI GPUs.

NV

Nvidia / Jensen Huang

Incumbent AI accelerator leader (~92% GPU market share) whose GB200/GB300 systems are the direct benchmark comparison; CEO publicly downplayed the chip as a competitive threat.

SE

SemiAnalysis

Independent semiconductor analysis firm that reviewed the InferenceX benchmark comparisons, calling the results industry-leading for a first-generation chip while flagging baseline caveats.

TS

TSMC

Foundry manufacturing Jalapeño's compute die on its 3nm-class N3P process.

CE

Celestica

Manufacturing and integration partner building board/rack-system integration and networking for the Jalapeño platform.

Fact Check

9 cited
  1. [1] Jalapeno in Nine Months: Did AI Just Break Chip Design Timelines?
  2. [2] OpenAI's Jalapeno Chip Is Outperforming Nvidia, AMD and Google Chips, SemiAnalysis Says
  3. [3] OpenAI Jalapeno ASIC at Hot Chips 2026
  4. [4] OpenAI Says Its Jalapeno Chip Beats Nvidia's GB300 in First Published Benchmarks
  5. [5] OpenAI Jalapeno: Better Than Nvidia
  6. [6] OpenAI's Jalapeno Chip Is Built for Fast Inference at Scale, Benchmarks Show
  7. [7] OpenAI Claims Its New Chips Can Outperform Nvidia Processors in Tests
  8. [8] Nvidia CEO Jensen Huang Responds to OpenAI's Jalapeno Chip
  9. [9] OpenAI's Jalapeno AI Chip and the New Threat to Nvidia

Source Articles

Top 5

THE SIGNAL.

Analysts

Presented Jalapeño's design philosophy and first benchmark results at Hot Chips 2026, characterizing the performance gains as a major advance over the state of the art.

Richard Ho
OpenAI's head/VP of Hardware

Found Jalapeño's efficiency and latency results credible and, unusually for a first-generation chip, industry-leading versus Nvidia, AMD and Google accelerators, but flagged that comparing it to GB300 rather than Nvidia's unreleased Rubin platform is an incomplete baseline.

SemiAnalysis (analyst team)
Independent semiconductor research firm

Downplayed Jalapeño as a competitive threat, framing custom-chip projects as common and often unsuccessful, and pointing to Nvidia's market share, revenue growth, and platform breadth as durable advantages.

Jensen Huang
CEO, Nvidia
The Crowd

We've designed and built our first AI chip: Jalapeño. Designed from the ground up by OpenAI and brought to production with @Broadcom, Jalapeño is purpose-built for the LLM workloads powering ChatGPT, Codex, the API, and future agentic products. Chips are foundational to the AI...

@@OpenAI22544

Since announcing Jalapeño, our first custom inference chip, we've been testing it and the system around it. The results show a major advance: more intelligence from every watt and faster responses, delivering both higher throughput and lower latency in one architecture without...

@@OpenAI14127

WOW!!! Jensen of $NVDA responds to OpenAI claims that their jalapeño chip is better than Nvidia's GB300. SHOTS FIRED!!! Reminds me of "You not talking to someone who woke up a loser" moment!! What a boss!!

@@RealNickMugalli1046

OpenAI Just Dropped Benchmarks for Their Own Chip, Jalapeño, and It's Beating Nvidia's GB300

@u/AskGpts283
Broadcast
OpenAI Says New Jalapeno Chips Outperformed Nvidia in Testing

OpenAI Says New Jalapeno Chips Outperformed Nvidia in Testing

OpenAI's Jalapeño Chip Could Change LLMS Forever

OpenAI's Jalapeño Chip Could Change LLMS Forever

OpenAI launched Jalapeno, their own AI chip!

OpenAI launched Jalapeno, their own AI chip!

OpenAI's Jalapeño custom inference chip debut — AI News | Agentic Brew