---
format: "aidr-story-markdown/v1"
id: "9fd3fcd02165c4af07f3f3d2c85dd236308376867f01cb16da48f54ac7971367"
canonical_url: "https://aidr.today/9fd3fcd0?lang=en"
title: "TPUv7 Outperforms Nvidia GB200 by 56% with First Kimi K3 Megakernel"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-23T20:43:36.000Z"
category: "Infra"
topics: ["tpu","nvidia","gb200","kimi","vllm","megakernel","elon-musk","chips"]
source_urls: ["https://marketbrief.now/ai/tpuv7-outperforms-nvidia-gb200-by-56percent-with-first-kimi-k3-megakerne-bac761a8","https://huggingnews.com/ai/tpuv7-outperforms-nvidia-gb200-by-56percent-with-first-kimi-k3-megakerne-bac761a8","https://x.com/elonmusk/status/2103329761690865846","https://x.com/SawyerMerritt/status/2103332409789776197","https://x.com/SawyerMerritt/status/2103335960092229948","https://huggingnews.com/ai/elon-musk-deploys-660000-nvidia-gb300-gpus-for-colossus-in-q4-d98b7ec6","https://marketbrief.now/ai/elon-musk-deploys-660000-nvidia-gb300-gpus-for-colossus-in-q4-d98b7ec6"]
summary: "vLLM maintainers and Inferact reported that TPUv7 hardware reached 709 tokens per second during low-concurrency decode for the Kimi K3 model. This performance beats the 450 tokens per second recorded by the Nvidia GB200 baseline for the same workload, representing a 56% increase in throughput. The improvement was driven by the implementation of a first-of-its-kind TPU megakernel optimization. The results are part of a broader trend toward the externalization of software for TPU hardware, moving these optimizations beyond internal Google environments. vLLM maintainers highlighted the benchmark to demonstrate how software-level changes to kernels can significantly alter the performance gap between competing AI accelerators."
---

# TPUv7 Outperforms Nvidia GB200 by 56% with First Kimi K3 Megakernel

> [Open the canonical story](<https://aidr.today/9fd3fcd0?lang=en>)

**Published:** 2026-09-23T20:43:36.000Z
**Category:** Infra
**Topics:** tpu, nvidia, gb200, kimi, vllm, megakernel, elon\-musk, chips

## Summary

vLLM maintainers and Inferact reported that TPUv7 hardware reached 709 tokens per second during low\-concurrency decode for the Kimi K3 model\. This performance beats the 450 tokens per second recorded by the Nvidia GB200 baseline for the same workload, representing a 56% increase in throughput\. The improvement was driven by the implementation of a first\-of\-its\-kind TPU megakernel optimization\. The results are part of a broader trend toward the externalization of software for TPU hardware, moving these optimizations beyond internal Google environments\. vLLM maintainers highlighted the benchmark to demonstrate how software\-level changes to kernels can significantly alter the performance gap between competing AI accelerators\.

## Sources

- [Story source](<https://marketbrief.now/ai/tpuv7-outperforms-nvidia-gb200-by-56percent-with-first-kimi-k3-megakerne-bac761a8>)
- [Story source](<https://huggingnews.com/ai/tpuv7-outperforms-nvidia-gb200-by-56percent-with-first-kimi-k3-megakerne-bac761a8>)
- [Story source](<https://x.com/elonmusk/status/2103329761690865846>)
- [Supporting source](<https://x.com/SawyerMerritt/status/2103332409789776197>)
- [Supporting source](<https://x.com/SawyerMerritt/status/2103335960092229948>)
- [Story source](<https://huggingnews.com/ai/elon-musk-deploys-660000-nvidia-gb300-gpus-for-colossus-in-q4-d98b7ec6>)
- [Story source](<https://marketbrief.now/ai/elon-musk-deploys-660000-nvidia-gb300-gpus-for-colossus-in-q4-d98b7ec6>)

