The San Francisco-based AI developer will begin installing its new "Jalapeño" inference chips into its compute infrastructure by the end of 2026. Across three public models on the InferenceX benchmark, the chip delivered 1.5 to 1.9 times more AI work per watt and a reduction in end-to-end latency of 1.7 to 3.6 times compared to Nvidia's GB300. Developed with Broadcom, the general purpose ASIC is designed specifically for running large language models rather than training them, offering 2.1 to 4.1 times higher performance on highly interactive workloads.

The chip is built on TSMC N3P with a 700W TDP and features 15.4TB/s of memory bandwidth, which is the highest for shipping or near shipping accelerators. OpenAI utilized its own Codex AI to assist in hardware design, reducing matrix engine area by 10% and completing the project from initial build to tape out in roughly 16 months. While the company is developing second and third generation versions, it noted that Jalapeño was not tested against Nvidia's newer Vera Rubin chips.

Sign in to suggest edits

Key sources

  1. SOURCE@openai“more intelligence from every watt and faster responses, delivering both higher throughput and lower latency in one architecture without sacrificing efficiency”x.com
  2. SOURCE@openai“Gen 2 is deep in development, and Gen 3 is taking shape”x.com
  3. SOURCE@cpmou2022“2.1–4.1× higher performance on highly interactive workloads”x.com
  4. SUPPORT@deitaone“Developed with Broadcom, Jalapeno targets AI inference and could lower data-center costs while handling large models”x.com
  5. SUPPORT@wallstengine“HBM4 with 15.4TB/s of memory bandwidth”x.com
  6. SUPPORT@openai“faster ChatGPT responses, more responsive Codex sessions and agents”x.com
  7. SOURCE@gdb“inference numbers published for jalapeno, team did an amazing job”x.com
  8. SUPPORT@edludlow“OAI's Richard Ho is on Bloomberg Tech today”x.com
Markdown