OpenAI's Jalapeño chip delivered 1.5x to 1.9x higher tokens per watt at peak throughput and 1.7x to 3.6x lower latency than Nvidia's GB200 and GB300 on GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T models. Developed with Broadcom, the hardware utilizes AI-generated assembly and kernels written in Gluon, a low-level language built on top of Triton that OpenAI engineers cannot reason about line by line.

The design approach removes human programmer usability as a primary constraint, relying instead on AI to write the compiler and optimize data movement across the hardware. OpenAI plans initial deployments before year end, with full scale operation starting in 2027.

Sign in to suggest edits

Key sources

  1. SOURCE@firesidealpha“OpenAI engineers scrolling their own kernel code had no idea what it did line by line”x.com
  2. SUPPORT@jukan05“AI may even begin designing architectures that humans would find awkward, unintuitive, or outright unpleasant to work with”x.com
  3. SOURCE@beth_kindig“1.5X to 1.9X higher tokens per watt at peak throughput and 1.7X to 3.6X lower latency”x.com
  4. SUPPORT@beth_kindig“initial deployments expected before year end”x.com
Markdown