Particle.news
Download on the App Store

OpenAI’s Jalapeño Chip Tops Public Benchmarks Against Nvidia Systems

OpenAI plans limited deployment of the HBM4-based Jalapeño in its data centers by the end of 2026.

Overview

  • OpenAI published InferenceX benchmark results on Tuesday showing Jalapeño delivered about 1.5–1.9x more AI work per watt at peak and roughly 1.7–3.6x lower end-to-end latency across GPT-OSS 120B, DeepSeek R1 and Kimi K2.5.
  • SemiAnalysis visited OpenAI’s labs to run the public InferenceX tests, but analysts warned the comparisons are not strictly like‑for‑like because Jalapeño uses HBM4 memory and was not tested against some of Nvidia’s newest HBM4 systems.
  • OpenAI says it used its own models to speed chip design and software: the team moved from initial design to tapeout in nine months and reported AI-generated kernel code ran about 1.5–1.8x faster on selected model blocks than human-written versions.
  • The company confirmed Jalapeño is focused on inference, will enter limited production use in OpenAI’s infrastructure by year-end, and that second- and third-generation designs are already in development while Nvidia and other partners remain in its fleet.
  • If real-world production volumes and pricing match lab claims, the chip could cut OpenAI’s operating cost per token and put pressure on inference margins across the AI supply chain, but independent, like‑for‑like tests and customer pricing remain open questions.