OpenAI posts first Jalapeño measurements: more work per watt than GB200/GB300
At Hot Chips 2026 OpenAI shared first results from its custom inference ASIC with Broadcom. SemiAnalysis InferenceX on GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T showed 1.5–1.9× more AI work per watt at peak throughput.

On August 25 OpenAI published the first measured results for Jalapeño, its custom inference ASIC built with Broadcom, at Hot Chips 2026. SemiAnalysis InferenceX compared the chip with Nvidia GB200/GB300 on GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T.
What we know
- Jalapeño is a custom inference ASIC built with Broadcom; first measured results were shown at Hot Chips 2026.
- On SemiAnalysis InferenceX, using GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T, it delivered 1.5–1.9× more AI work per watt at peak throughput.
- End-to-end latency was 1.7–3.6× lower, and interactive performance was 2.1–4.1× versus the Nvidia GB200/GB300 comparison.
- Richard Ho said a small-volume deployment comes at the end of 2026, with higher volume in 2027.
Takeaways
- These are first lab measurements, not a shipping volume SKU.
- The comparison set is GB200/GB300 on three named large models.
- Ho’s calendar is late-2026 for a small run and 2027 for higher volume.
Source: OpenAI


