OpenAI Unveils Jalapeño, Its First Custom AI Inference Chip
The chip aims to lower inference costs, reducing reliance on outside GPU suppliers.
Overview
- OpenAI and Broadcom publicly revealed the Jalapeño on Wednesday, June 24, 2026, and said working lab samples are running the GPT-5.3-Codex-Spark model with expected power and performance.
- Broadcom's CEO Hock Tan reported preliminary lab tests show roughly 50% cost savings versus typical AI GPUs and claimed efficiency comparable to other leading accelerators, though those figures come from company testing rather than independent benchmarks.
- OpenAI says engineers finished the design in about nine months and sent the chips to TSMC for fabrication, with Celestica contracted to build the server systems that will house the chips.
- OpenAI plans to begin integrating final Jalapeño chips into Microsoft and partner data centers by the end of 2026 and describes the chip as the first in a multi‑generation roadmap.
- The move reflects a wider industry shift toward bespoke accelerators to cut operating costs and gain control over compute; the rollout will hinge on high‑bandwidth memory suppliers and a Broadcom financing vehicle backed by Apollo and Blackstone to support large chip purchases.