OpenAI unveils Jalapeño inference chip, claims it beats Nvidia
TECH

OpenAI unveils Jalapeño inference chip, claims it beats Nvidia

43+
Signals

Strategic Overview

  • 01.
    OpenAI presented the first live benchmark results for its custom inference chip, Jalapeño, at the Hot Chips conference at Stanford, claiming it outperforms Nvidia's GB300 on throughput-per-watt and response latency.
  • 02.
    Jalapeño was co-developed with Broadcom (silicon and networking) and Celestica (systems integration), stemming from an October 2025 deal to co-develop 10 gigawatts of custom AI accelerators.
  • 03.
    On SemiAnalysis's public InferenceX benchmark suite, Jalapeño delivered 1.5x-1.9x more throughput per kilowatt and 1.7x-3.6x lower end-to-end latency than Nvidia's GB200 and GB300 rack systems.
  • 04.
    OpenAI plans to begin deploying Jalapeño in its own data centers in limited volume by the end of 2026, with a more significant rollout in 2027.

Deep Analysis

The Fine Print Behind the Benchmark Chart

The headline numbers come from SemiAnalysis's public InferenceX suite, but the fine print matters. As the firm itself notes: "All numbers are provided to us by OpenAI. We verified the InferenceX runs in person in the lab, but we did not run the full suite of InferenceX benchmarks" [1]. The comparison also excludes Nvidia's newest Vera Rubin systems, which recently began shipping [2]- meaning Jalapeño was measured against the outgoing GB200/GB300 generation rather than Nvidia's current answer. No multi-turn or long-context (AgentX) benchmarks were run, only single-turn 8k/1k tests [1]. Online technical discussion has also raised a further nuance worth flagging as unverified: that the Jalapeño runs reportedly skipped speculative decoding, a technique commonly used in Rubin-class comparisons, which some argue could flatter Jalapeño's relative numbers - a claim from community discussion, not confirmed in the underlying reporting, so it should be read as an open question rather than a settled fact.

Betting on the Full Stack

OpenAI's head of hardware, Richard Ho, frames the project simply: "Jalapeño can serve more AI work per unit of power, while also returning responses more quickly" [3]. The architecture behind that claim is deliberate - OpenAI says it built Jalapeño "to minimize data movement and communication delays" [4], keeping model state and KV cache close via a large SRAM cache paired with 216 GB of HBM4 memory running at 15.4 TB/s per chip [5]. That design targets the prefill and communication phases that typically bottleneck inference at scale - the layer OpenAI can only optimize by controlling silicon, networking, and models together, rather than buying commodity GPU systems off the shelf. The project also moved unusually fast: from architecture concept in late 2024 to RTL freeze and tapeout in 2025, with Codex and then ChatGPT running on the chip in early 2026 [5], using OpenAI's own AI models paired with the XLS hardware description language to accelerate development [5]- a roughly nine-month RTL-to-tapeout cycle that is fast even by custom-silicon standards [6]- a pace that drew attention beyond OpenAI's own materials, with online commentary singling it out as the single most striking detail in the whole announcement. In a Bloomberg Tech interview, Ho put the stakes in blunter terms: achieving high throughput and low latency at the same time, on the same chip, is an industry first, he said, and it's the kind of gain that only comes from controlling the full stack - models, software, and silicon together - rather than assembling systems from off-the-shelf Nvidia parts. He also said OpenAI deliberately chose SemiAnalysis's InferenceX suite because it is open-source and neutral rather than a benchmark OpenAI itself controls.

Where Nvidia's Moat Still Holds

Jalapeño is an inference-only chip - it doesn't train models, the workload where Nvidia's hardware remains unchallenged [2]. That distinction caps how much independence OpenAI can realistically claim from a single announcement. Even under an aggressive rollout - limited volume by the end of 2026, a more significant deployment in 2027 [7]- OpenAI's training stack, and the billions of dollars in GPU commitments behind it, stays tied to Nvidia. Jalapeño is a claim about the economics of serving models cheaply at scale, not about building them.

Nvidia's Margin Question, and Who's Skeptical

CNBC frames Jalapeño as a fresh threat to Nvidia's margins as custom silicon from model makers and hyperscalers takes share of the fast-growing inference market [8]. Not everyone buys the framing, though: analyst Daniel Newman credited Broadcom for Jalapeño's engineering while calling 'Nvidia killer' talk exhausting hype [9]. Reaction online split along similar lines - some treated the results skeptically, as a vendor-run demo to be trusted only once independently reproduced, while others pointed to SemiAnalysis's in-lab verification as reason to take the numbers seriously. The disagreement is less about the raw figures than about how much weight a self-reported, first-generation chip announcement should carry against years of Nvidia's production track record. The public framing has also shifted markedly over time: when Broadcom CEO Hock Tan first showed Jalapeño to Sam Altman and Greg Brockman back in June 2026, he told Reuters its performance was 'apparently the same' as Nvidia's latest Blackwell chips and Google's TPUs - a far more modest claim than the outperformance numbers OpenAI presented at Hot Chips two months later. That gap between the original, cautious framing and August's benchmark-backed claims is itself part of the story: the narrative around Jalapeño escalated considerably between unveiling and proof.

Historical Context

2024
Jalapeño's architecture concept phase began.
2025-10
OpenAI and Broadcom struck a deal to co-develop 10 gigawatts of custom AI accelerators, the origin of the Jalapeño program.
2025
RTL freeze occurred during 2025, followed by tapeout in late 2025 (sent to TSMC for manufacturing); design-to-tapeout took roughly nine months.
2026-06
OpenAI first unveiled Jalapeño alongside Broadcom; a working sample was handed to Sam Altman and Greg Brockman by Broadcom CEO Hock Tan; early samples reportedly showed cost savings of roughly 50% versus typical AI GPUs.
2026-08-25
OpenAI presented first live Jalapeño benchmark results at the Hot Chips conference at Stanford University, claiming performance and efficiency advantages over Nvidia's GB200 and GB300 systems.

Power Map

Key Players
Subject

OpenAI unveils Jalapeño inference chip, claims it beats Nvidia

OP

OpenAI

Designed Jalapeño around its own LLM roadmap, presented benchmark results at Hot Chips, and plans to deploy the chip in its own data centers starting late 2026.

BR

Broadcom

Co-developed the chip's silicon and networking, including the Tomahawk6 switches used in the system fabric, as part of an October 2025 deal to co-develop 10GW of custom AI accelerators with OpenAI.

CE

Celestica

Handled systems integration for Jalapeño racks.

NV

Nvidia

Incumbent GPU maker whose GB200, GB300, and Blackwell-class systems form the comparison baseline; faces potential margin pressure as custom silicon gains ground in inference.

SE

SemiAnalysis

Independent analyst firm whose public InferenceX benchmark suite was used for the tests; verified the runs in person in the lab but did not run the full benchmark suite itself.

TS

TSMC

Manufactures the Jalapeño chip.

Fact Check

9 cited
  1. [1] OpenAI Jalapeno: Better Than Nvidia
  2. [2] OpenAI Jalapeno Chip Nvidia Benchmark Results
  3. [3] OpenAI's Jalapeno Chip Is Built for Fast Inference at Scale, Benchmarks Show
  4. [4] OpenAI's Upcoming Jalapeno Chip Looks Like It'll Be an Inference Beast
  5. [5] OpenAI Jalapeno ASIC at Hot Chips 2026
  6. [6] OpenAI Claims AI Chip Breakthrough With Broadcom Collaboration: Jalapeno 'Beats' Nvidia's GB300 in Tests
  7. [7] OpenAI Says Its Jalapeno Chip Beats Nvidia's GB300 in First Published Benchmarks
  8. [8] OpenAI's Jalapeno AI Chip Poses a Fresh Threat to Nvidia
  9. [9] Daniel Newman Says Broadcom Deserves Credit for OpenAI's Jalapeno Chip But Calls 'Nvidia Killer' Hype Exhausting

Source Articles

Top 5

THE SIGNAL.

Analysts

Says Jalapeño beats Blackwell across almost all tested scenarios without being tuned for a single point on the performance curve, excelling at both low-latency and high-throughput ends, but flags that its own numbers came from OpenAI and that no multi-turn, long-context benchmarks were run.

SemiAnalysis
Independent semiconductor analyst firm

Frames Jalapeño's value proposition as doing more AI work per unit of power while responding faster to users.

Richard Ho
Head of Hardware, OpenAI

Argues Broadcom deserves credit for Jalapeño's engineering and calls the framing of the chip as an 'Nvidia killer' exhausting hype.

Daniel Newman
Industry analyst
The Crowd

Since announcing Jalapeño, our first custom inference chip, we've been testing it and the system around it. The results show a major advance: more intelligence from every watt and faster responses, delivering both higher throughput and lower latency in one architecture without...

@@OpenAI13143

OpenAI says its new Broadcom-built Jalapeno AI chip outperformed $NVDA GB300 in both throughput per watt and response latency in its testing. The chip is designed specifically for inference, not model training, and runs at around 700 watts. OpenAI plans to begin deploying it for...

@@wallstengine415

OPENAI TOOK ITS FIRST AI CHIP FROM DESIGN TO TAPEOUT IN NINE MONTHS. Now it says the chip beats systems built on NVIDIA's GB200 and GB300 on the curve agents care about. Today OpenAI published the first measured results from Jalapeño, its custom inference chip. Across GPT-OSS...

@@lagerskoy21

OpenAI Just Dropped Benchmarks for Their Own Chip, Jalapeño, and It's Beating Nvidia's GB300

@AskGpts250
Broadcast
OpenAI Says New Jalapeno Chips Outperformed Nvidia in Testing

OpenAI Says New Jalapeno Chips Outperformed Nvidia in Testing

OpenAI's Jalapeño Chip Could Change LLMS Forever

OpenAI's Jalapeño Chip Could Change LLMS Forever

OpenAI's Jalapeño: The Chip That Ends the GPU Hype?

OpenAI's Jalapeño: The Chip That Ends the GPU Hype?