OpenAI Launches Its Own Chip: Jalapeño — Built for LLM Inference, GPU Dominance Era Ends
OpenAI and Broadcom jointly unveiled Jalapeño, OpenAI's first custom inference chip designed from scratch for LLM workloads. It taped out in just 9 months and is already running GPT-5.3-Codex-Spark in the lab.
💡 What You Will Learn
OpenAI and Broadcom jointly unveiled Jalapeño, OpenAI's first custom inference chip designed from scratch for LLM workloads. It taped out in just 9 months and is already running GPT-5.3-Codex-Spark in
OpenAI finally made its move — not relying on NVIDIA, but building its own chip. Yesterday, OpenAI and Broadcom jointly unveiled Jalapeño, OpenAI's first custom inference chip. Not a modified IP, but designed from scratch. 9 months from design to tape-out. Engineering samples are already running GPT-5.3-Codex-Spark in the lab.
Jalapeño is not a 'modified GPU.' Existing AI chips mostly derive from GPU architecture — originally built for graphics rendering, re-purposed for LLMs. Jalapeño started from scratch, redefining chip architecture from LLM inference requirements. The core idea: reduce data movement. Balance compute, memory, and network resources so real utilization approaches theoretical peak.
The result: per-watt performance already substantially exceeds current state-of-the-art chips. OpenAI's official statement: 'substantially better than current state-of-the-art.' And it's not a one-model special — it's designed to efficiently run all LLMs, not just GPT series.
9 months is the fastest ASIC development cycle in history. Industry standard for a high-performance AI chip is 18-24 months. Jalapeño cut it in half. Two reasons: (1) 'The people designing the chip are the ones running it every day' — OpenAI knows every layer of ChatGPT, Codex, and API workloads perfectly. (2) OpenAI used its own models to accelerate the chip design process — the models you use in ChatGPT are helping design the hardware that runs them.
Jalapeño is not a one-off. It's the first generation of a computing platform series. Deployment starts late 2026, in partnership with Microsoft building gigawatt-scale data centers. Broadcom provides Tomahawk networking chips for large-scale interconnect, Celestica handles board and rack system integration.
For ordinary users, the direct impact is one word: cheaper. AI inference costs are dominated by chips. Custom chips eliminate NVIDIA's margin layer, plus the massive per-watt improvement means inference costs will inevitably drop. Brockman's own words: 'Make computing more abundant, make AI faster, more reliable, and cheaper.'
Personal view: OpenAI isn't 'competing with NVIDIA' — it's consolidating product, model, infrastructure, and chip layers under one roof. These four layers previously had massive information loss. Now they talk in the same building. Faster iteration, lower costs, stronger competitive moat.
For NVIDIA, this is bad news. H100/B200 customers are building their own chips — Google has TPU, Amazon has Trainium, Microsoft has Maia, Meta has MTIA. The GPU monopoly era for AI is over.
But the bigger story: when building chips becomes faster and cheaper, the barrier to entry for AI drops further. 9-month tape-out shows that AI is reshaping the semiconductor industry itself.
Written by our editorial team; tools listed here are tested or verified against public sources. Links point to official sites or GitHub repos for reference only — no paid placements.
