OpenAI has provided further details about its custom AI accelerator 'Jalapeño,' developed in partnership with Broadcom.

At the Hot Chips event, the company revealed that the processor is rated at 700 watts, but the chip's measured sustained power remained at or below 550W on the workloads tested.

OpenAI Jalapeño
– OpenAI

OpenAI will deploy Jalapeño in racks featuring 128 chips, while "a full pod is 2,048 ASICs," the company's VP of hardware, Richard Ho, said in a presentation attended by DCD. The 128 deployment is capable of 1.7 exaflops of 4-bit compute, with 27.5TB of HBM4. Each package has 15.4TBps of memory bandwidth.

Ho said that the company tested the chip on three models, its own GPT-OSS-120B, as well as DeepSeek's R1 and Moonshot AI's Kimi K2.5. "Across the three models we tried this on, we showed that the architecture works across not just internal models, but others," Ho said. "This is a general purpose, very flexible accelerator."

On SemiAnalysis’ InferenceX benchmarks, the company claimed that Jalapeño-based systems provided between 1.5x and 1.9x more “AI work” at peak throughput, and 1.7x to 3.6x lower end-to-end latency than rival platforms on the three models. On ultra-low latency, it performed 2.1x to 4.1x faster.

While the company claims it bests the inference performance of Nvidia's GB200 NVL72 and GB300 NVL72 rack systems, OpenAI does not plan to slow down the deployment of rival hardware. "This is part of our overall compute strategy, which includes our very, very good partners at Nvidia, Cerebras, and others," Ho said. "We have a lot of compute needs, and we're not going to get them with one shot."

Similarly, when asked about whether OpenAI would sell its chips to other companies, Ho said: "We have a lot of compute needs within OpenAI, it's going to take a lot of compute to meet those, we are struggling to have enough."

OpenAI claims that Jalapeño was developed in part with AI, making the process faster than ever. As both the chips and the AI improve, this is expected to get ever better. "We were able to fit more compute into the chip using AI," Ho said, adding that the Gen 2 version of the chip is "deep into development" and the third is underway.

Jalapeño is expected to be deployed in limited quantities later this year, and more widely deployed next year.