The AI Designed Its Own Chip
OpenAI's chip effort moved from initial design to tapeout in an estimated nine to sixteen months [5], an unusually fast timeline that one industry analysis flagged as a possible record for chip design [1]- compressed largely because OpenAI turned its own models loose on the hardware itself. AI-written kernel implementations for Jalapeño outperformed human expert kernels by 1.5 to 1.8 times [2], with AI also used to accelerate verification loops and low-level kernel optimization throughout the design process [2]. A technical teardown found the AI-tuned matrix units delivered roughly a 56% gain on BF16 multiply operations while shrinking matrix-unit die area by about 10% [3]. In a Bloomberg Technology interview, Richard Ho pointed to a different, faster-moving marker of that speed: three very different AI models were brought up and running on the new hardware within a couple of months of OpenAI receiving it, which he cited as evidence that the programming model itself had matured fast enough to keep pace with the silicon. The resulting single chip packs 13.4 PFLOPS of MXFP4 compute, 216 GiB of HBM4 memory, and 15.4 TB/s of bandwidth in a 700W envelope [3], scaling to 27 EFLOPS and 432 TiB of memory across a 2,048-chip pod [3]. The notable part isn't just that a lab built a chip - it's that the chip's own creation leaned on the AI it exists to run, folding model-assisted engineering into a discipline that historically took human teams years.



