Broadcom-built silicon from OpenAI provides 1.7 to 3.6 times lower end-to-end latency than Nvidia's GB300 system in new performance benchmarks. The chip, named Jalapeño, delivers between 1.5 and 1.9 times more AI work per watt and is designed exclusively for inference rather than model training. OpenAI will begin deploying the hardware within its compute infrastructure by the end of 2026, with gains extending across models such as GPT-OSS, DeepSeek R1, and Kimi K2.5.

Built on TSMC N3P with 15.4TB/s of memory bandwidth, the 700W chip was developed using AI-assisted design via Codex to shorten the timeline from initial design to tape-out. Current benchmarks use A0 silicon, while a B0 update in production is slated to improve power efficiency by 25%. OpenAI is targeting an initial deployment of around 100MW and has a second-generation chip already approaching tape-out.

Sign in to suggest edits

Key sources

  1. SOURCE@gdb“inference numbers published for jalapeno, team did an amazing job”x.com
  2. SUPPORT@edludlow“OAI's Richard Ho is on Bloomberg Tech today”x.com
  3. SUPPORT@cpmou2022“2.1–4.1× higher performance on highly interactive workloads”x.com
  4. SUPPORT@wallstengine“OpenAI says a second-generation chip is already approaching tape-out, while work on a third generation has begun.”x.com
  5. SUPPORT@wallstengine“HBM4 with 15.4TB/s of memory bandwidth”x.com
  6. SUPPORT@deitaone“However, Jalapeno wasn’t tested against Nvidia’s newer Vera Rubin chips.”x.com
  7. SUPPORT@stocksavvyshay“Jalapeno reportedly delivered more AI work per watt and faster response times as inference becomes a larger share of AI compute.”x.com
Markdown