OpenAI’s 700W Jalapeño ASIC outpaces 1,400W Nvidia flagship GPU — claims up to 1.9x throughput per kilowatt and 3.6x lower latency, co-developed with Broadcom

OpenAI’s 700W Jalapeño ASIC outpaces 1,400W Nvidia flagship GPU — claims up to 1.9x throughput per kilowatt and 3.6x lower latency, co-developed with Broadcom

When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works .

The tests covered three open models: GPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI's 1-trillion-parameter Kimi K2.5, with OpenAI reporting its widest leads at low-latency operating points, where it claims 8.6 times to 104.3 times more throughput per kilowatt at the GB300's fastest previous time-between-tokens settings.

OpenAI normalized the results to each accelerator's published package TDP, though it said Jalapeño's measured sustained power stayed at or below 550W in testing. An appendix comparison using all-in utility power per accelerator, 1.18kW for Jalapeño against 2.55kW for the GB300, produces narrower gaps, as does pitting Jalapeño against a GB300 running multi-token prediction, where the peak efficiency lead shrinks to roughly 1.5 times.

Jalapeño wasn't tested against Vera Rubin, the Nvidia platform that's slated to power the first gigawatt of Nvidia systems OpenAI agreed to deploy in the second half of 2026. The chip also doesn't train models, the workload where Nvidia's hardware remains unchallenged. In addition, the major comparisons ran Jalapeño's single-token prediction against GB300 configurations doing the same, even though Nvidia deployments commonly use multi-token prediction in production. SemiAnalysis , which said it ran InferenceX with OpenAI engineers in the company's lab, described the part as "beating every Nvidia, AMD, and Google chip we have been able to test."

Each Jalapeño package, unveiled in June after a nine-month RTL-to-tapeout cycle, pairs its compute die with six HBM4 stacks, totaling 216 GiB at 15.4 TB/s. The GB300 carries 288GB of HBM3E at a 1,400W rating, so per watt of rated power, OpenAI's chip packs roughly 50% more memory. The company's Hot Chips presentation states that the main bottleneck its architecture targets is exposing aggregate HBM bandwidth, not adding more of it.

Broadcom and OpenAI unveil custom-built Jalapeño inference processor

Key considerations

  • Investor positioning can change fast
  • Volatility remains possible near catalysts
  • Macro rates and liquidity can dominate flows

Reference reading

More on this site

Informational only. No financial advice. Do your own research.

Leave a Comment