Hardware

OpenAI Unveils Jalapeño, Its In-House Inference Chip Developed with Broadcom

Early benchmarks claim substantial efficiency and latency gains over Nvidia’s flagship systems ahead of a late-2026 rollout.

  • OpenAI has disclosed the first benchmark metrics for Jalapeño, a proprietary inference accelerator engineered in collaboration with Broadcom, marking the AI lab's most concrete step yet toward esta…
  • According to data published by OpenAI using the public InferenceX benchmark, the processor delivered noticeable performance gains across large-scale architectures, including GPT-OSS 120B, DeepSeek…
  • The benchmark comparisons show the sharpest divergence in real-time response times.
OpenAI Unveils Jalapeño, Its In-House Inference Chip Developed with BroadcomThe Scale Report

OpenAI has disclosed the first benchmark metrics for Jalapeño, a proprietary inference accelerator engineered in collaboration with Broadcom, marking the AI lab's most concrete step yet toward establishing its own custom silicon supply.

According to data published by OpenAI using the public InferenceX benchmark, the processor delivered noticeable performance gains across large-scale architectures, including GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. When measured against Nvidia’s GB200 and GB300 systems, Jalapeño achieved 1.5x to 1.9x more AI work per watt.

Challenging Nvidia on Latency

The benchmark comparisons show the sharpest divergence in real-time response times. OpenAI claimed Jalapeño cut end-to-end latency by 1.7x to 3.6x relative to Nvidia's hardware, while achieving up to a 4.1x performance increase on highly interactive tasks where rapid token delivery is critical.

OpenAI plans to begin deploying Jalapeño in its production clusters in late 2026. The organization indicated that subsequent generations of the architecture are already in active development, even as it plans to maintain a hybrid operational model that keeps Nvidia silicon in rotation for relevant workloads.

The Silicon Diversification Play

Designing custom application-specific integrated circuits (ASICs) follows the playbook established by major cloud operators such as Google with its TPUs and Amazon with Trainium. As user queries scale and ongoing inference costs eclipse initial model training expenditures, reliance on general-purpose GPUs introduces margin compression that in-house hardware is uniquely positioned to offset.

Reporting based on coverage from @technology on Instagram.

Read next