---
format: "aidr-story-markdown/v1"
id: "6c0f4a8ded77bbab9e50fc6e789b55ec72a2bc48913120e860d27ebe50a751f8"
canonical_url: "https://aidr.today/6c0f4a8d?lang=en"
title: "OpenAI Jalapeño Delivers 1.9x More Work Per Watt in First Custom Inference Chip"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-25T16:48:16.000Z"
category: "Chips"
topics: ["openai","nvidia","inference","chips"]
source_urls: ["https://huggingnews.com/ai/openai-jalapeno-delivers-19x-more-work-per-watt-in-first-custom-inferenc-cd2b52b1","https://x.com/gdb/status/2092273740239552780","https://x.com/EdLudlow/status/2092251479793184938","https://x.com/CPMou2022/status/2092261275967307803","https://x.com/wallstengine/status/2092252412711215158","https://x.com/wallstengine/status/2092254221458665537","https://x.com/DeItaone/status/2092251884895527366","https://x.com/StockSavvyShay/status/2092252634057236891"]
summary: "Broadcom-built silicon from OpenAI provides 1.7 to 3.6 times lower end-to-end latency than Nvidia's GB300 system in new performance benchmarks. The chip, named Jalapeño, delivers between 1.5 and 1.9 times more AI work per watt and is designed exclusively for inference rather than model training. OpenAI will begin deploying the hardware within its compute infrastructure by the end of 2026, with gains extending across models such as GPT-OSS, DeepSeek R1, and Kimi K2.5. Built on TSMC N3P with 15.4TB/s of memory bandwidth, the 700W chip was developed using AI-assisted design via Codex to shorten the timeline from initial design to tape-out. Current benchmarks use A0 silicon, while a B0 update in production is slated to improve power efficiency by 25%. OpenAI is targeting an initial deployment of around 100MW and has a second-generation chip already approaching tape-out."
---

# OpenAI Jalapeño Delivers 1\.9x More Work Per Watt in First Custom Inference Chip

> [Open the canonical story](<https://aidr.today/6c0f4a8d?lang=en>)

**Published:** 2026-08-25T16:48:16.000Z
**Category:** Chips
**Topics:** openai, nvidia, inference, chips

## Summary

Broadcom\-built silicon from OpenAI provides 1\.7 to 3\.6 times lower end\-to\-end latency than Nvidia's GB300 system in new performance benchmarks\. The chip, named Jalapeño, delivers between 1\.5 and 1\.9 times more AI work per watt and is designed exclusively for inference rather than model training\. OpenAI will begin deploying the hardware within its compute infrastructure by the end of 2026, with gains extending across models such as GPT\-OSS, DeepSeek R1, and Kimi K2\.5\. Built on TSMC N3P with 15\.4TB/s of memory bandwidth, the 700W chip was developed using AI\-assisted design via Codex to shorten the timeline from initial design to tape\-out\. Current benchmarks use A0 silicon, while a B0 update in production is slated to improve power efficiency by 25%\. OpenAI is targeting an initial deployment of around 100MW and has a second\-generation chip already approaching tape\-out\.

## Sources

- [Story source](<https://huggingnews.com/ai/openai-jalapeno-delivers-19x-more-work-per-watt-in-first-custom-inferenc-cd2b52b1>)
- [Story source](<https://x.com/gdb/status/2092273740239552780>)
- [Supporting source](<https://x.com/EdLudlow/status/2092251479793184938>)
- [Supporting source](<https://x.com/CPMou2022/status/2092261275967307803>)
- [Supporting source](<https://x.com/wallstengine/status/2092252412711215158>)
- [Supporting source](<https://x.com/wallstengine/status/2092254221458665537>)
- [Supporting source](<https://x.com/DeItaone/status/2092251884895527366>)
- [Supporting source](<https://x.com/StockSavvyShay/status/2092252634057236891>)

