Overview
- OpenAI published SemiAnalysis InferenceX results at the Hot Chips conference on Tuesday showing Jalapeño delivered about 1.5–1.9 times more AI work per watt and 1.7–3.6 times lower end-to-end latency across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T.
- The tests measured full inference flows on a public benchmark and used published competitor power figures for comparison, but they did not include Nvidia’s newest Rubin/Vera Rubin generation, so the results are not a direct like‑for‑like comparison with Nvidia’s latest HBM4 systems.
- OpenAI says Jalapeño is an inference-focused ASIC built with Broadcom and Celestica that rates 700 watts TDP yet sustained roughly 550 watts in tests, and the company plans small‑volume deployment by the end of 2026 with a wider ramp through 2027.
- OpenAI accelerated chip development with its own models, moving from design to tapeout in nine months, and reports AI-generated implementations ran 1.5–1.8 times faster on selected attention and mixture‑of‑experts blocks than human-written versions.
- Analysts say hyperscaler-designed chips like Jalapeño could pressure Nvidia’s inference economics over time and help lower the cost and latency of multi-step AI workflows, while Nvidia and other partners remain important for training and broader compute needs.