vLLM maintainers and Inferact reported that TPUv7 hardware reached 709 tokens per second during low-concurrency decode for the Kimi K3 model. This performance beats the 450 tokens per second recorded by the Nvidia GB200 baseline for the same workload, representing a 56% increase in throughput. The improvement was driven by the implementation of a first-of-its-kind TPU megakernel optimization. The results are part of a broader trend toward the externalization of software for TPU hardware, moving these optimizations beyond internal Google environments. vLLM maintainers highlighted the benchmark to demonstrate how software-level changes to kernels can significantly alter the performance gap between competing AI accelerators.

Sign in to suggest edits

Key sources

  1. SOURCEmarketbrief.now
  2. SOURCEhuggingnewshuggingnews.com
  3. SOURCE@elonmusk“Another 220k GB300 will be fully operational next week and another 220k in November. If we get lucky, yet another 220k GB300 by late December.”x.com
  4. SUPPORT@sawyermerritt“Next week, @SpaceX will be bringing online 220,000 Nvidia GB300 GPU (worth $15+ billion)”x.com
  5. SUPPORT@sawyermerritt“That's over $100 billion of GPUs combined, and brought online in record time. Over 2 GW of compute.”x.com
  6. SOURCEhuggingnewshuggingnews.com
  7. SOURCEmarketbrief.now
Markdown