OpenAI has unveiled details about its Jalapeño chip, an Application-Specific Integrated Circuit (ASIC) developed in partnership with Broadcom, designed specifically for AI inference—the process of running trained AI models to complete tasks or deploy agents. According to OpenAI hardware vice president Richard Ho, Jalapeño offers the "best of both worlds" by combining lower latency and higher throughput, addressing a typical trade-off in AI systems.

In benchmark tests using InferenceX, Jalapeño outperformed Nvidia's GB200 and GB300 superchips across three AI models. The chip delivered 1.5 to 1.9 times more AI work per watt across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T, while achieving 1.7 to 3.6 times lower end-to-end latency across the same models.

OpenAI plans to deploy Jalapeño in "small volumes" by the end of this year, ramping up production into 2027. However, the company does not expect to replace its entire chip lineup with Jalapeño and will continue working with partners like Nvidia. OpenAI also indicated plans to develop second and third generations of the chip.