---
format: "aidr-story-markdown/v1"
id: "18be44edb94925fb59eb70eb6c8225a9c57c06aae38819aa7c0cc6e73a6be013"
canonical_url: "https://aidr.today/18be44ed?lang=en"
title: "OpenAI's First Custom AI Chip Beats Nvidia GB300 by 1.9x on Power Efficiency"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-25T18:52:21.000Z"
category: "Chips"
topics: ["openai","nvidia","chips","inference","infra"]
source_urls: ["https://huggingnews.com/ai/openais-first-custom-ai-chip-beats-nvidia-gb300-by-19x-on-power-efficien-aa9adaf6","https://x.com/OpenAI/status/2092300846675505602","https://x.com/OpenAI/status/2092300851482108064","https://x.com/CPMou2022/status/2092261275967307803","https://x.com/DeItaone/status/2092251884895527366","https://x.com/wallstengine/status/2092254221458665537","https://x.com/OpenAI/status/2092300848483143735","https://x.com/gdb/status/2092273740239552780"]
summary: "The San Francisco-based AI developer will begin installing its new \"Jalapeño\" inference chips into its compute infrastructure by the end of 2026. Across three public models on the InferenceX benchmark, the chip delivered 1.5 to 1.9 times more AI work per watt and a reduction in end-to-end latency of 1.7 to 3.6 times compared to Nvidia's GB300. Developed with Broadcom, the general purpose ASIC is designed specifically for running large language models rather than training them, offering 2.1 to 4.1 times higher performance on highly interactive workloads. The chip is built on TSMC N3P with a 700W TDP and features 15.4TB/s of memory bandwidth, which is the highest for shipping or near shipping accelerators. OpenAI utilized its own Codex AI to assist in hardware design, reducing matrix engine area by 10% and completing the project from initial build to tape out in roughly 16 months. While the company is developing second and third generation versions, it noted that Jalapeño was not tested against Nvidia's newer Vera Rubin chips."
---

# OpenAI's First Custom AI Chip Beats Nvidia GB300 by 1\.9x on Power Efficiency

> [Open the canonical story](<https://aidr.today/18be44ed?lang=en>)

**Published:** 2026-08-25T18:52:21.000Z
**Category:** Chips
**Topics:** openai, nvidia, chips, inference, infra

## Summary

The San Francisco\-based AI developer will begin installing its new "Jalapeño" inference chips into its compute infrastructure by the end of 2026\. Across three public models on the InferenceX benchmark, the chip delivered 1\.5 to 1\.9 times more AI work per watt and a reduction in end\-to\-end latency of 1\.7 to 3\.6 times compared to Nvidia's GB300\. Developed with Broadcom, the general purpose ASIC is designed specifically for running large language models rather than training them, offering 2\.1 to 4\.1 times higher performance on highly interactive workloads\. The chip is built on TSMC N3P with a 700W TDP and features 15\.4TB/s of memory bandwidth, which is the highest for shipping or near shipping accelerators\. OpenAI utilized its own Codex AI to assist in hardware design, reducing matrix engine area by 10% and completing the project from initial build to tape out in roughly 16 months\. While the company is developing second and third generation versions, it noted that Jalapeño was not tested against Nvidia's newer Vera Rubin chips\.

## Sources

- [Story source](<https://huggingnews.com/ai/openais-first-custom-ai-chip-beats-nvidia-gb300-by-19x-on-power-efficien-aa9adaf6>)
- [Story source](<https://x.com/OpenAI/status/2092300846675505602>)
- [Story source](<https://x.com/OpenAI/status/2092300851482108064>)
- [Story source](<https://x.com/CPMou2022/status/2092261275967307803>)
- [Supporting source](<https://x.com/DeItaone/status/2092251884895527366>)
- [Supporting source](<https://x.com/wallstengine/status/2092254221458665537>)
- [Supporting source](<https://x.com/OpenAI/status/2092300848483143735>)
- [Story source](<https://x.com/gdb/status/2092273740239552780>)

