---
format: "aidr-story-markdown/v1"
id: "7ef96a45f8c78a770275812e944d5c42c383829ef5935ef26939412195d902d3"
canonical_url: "https://aidr.today/7ef96a45?lang=en"
title: "GLM-5.3 Hits 21.4x PyTorch Speedup on KernelBench-Mega, Doubling GLM-5.2 Performance"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-22T14:48:07.000Z"
category: "Research"
topics: ["inference","chips","llm"]
source_urls: ["https://huggingnews.com/ai/glm-53-hits-214x-pytorch-speedup-on-kernelbench-mega-doubling-glm-52-per-930bfc4c","https://x.com/elliotarledge/status/2091125950243033209","https://x.com/Zai_org/status/2091141282714587232"]
summary: "Zai's latest model utilized Kimi-Linear Decode on an RTX PRO 6000 GPU to optimize inference throughput. The GLM-5.3 version reached 21.4x the PyTorch baseline on the KernelBench-Mega benchmark, nearly doubling the 11.1x speedup of GLM-5.2. This iteration also passed the single-launch gate, a technical requirement that the previous version failed. The performance gain relies on a CUDA `load_inline` launch with persistent 512-thread CTAs and 30 grid barriers. The implementation keeps Int4 dequant inside the GEMV and absorbs Multi-Head Latent Attention (MLA) so the fat KV is not materialized."
---

# GLM\-5\.3 Hits 21\.4x PyTorch Speedup on KernelBench\-Mega, Doubling GLM\-5\.2 Performance

> [Open the canonical story](<https://aidr.today/7ef96a45?lang=en>)

**Published:** 2026-08-22T14:48:07.000Z
**Category:** Research
**Topics:** inference, chips, llm

## Summary

Zai's latest model utilized Kimi\-Linear Decode on an RTX PRO 6000 GPU to optimize inference throughput\. The GLM\-5\.3 version reached 21\.4x the PyTorch baseline on the KernelBench\-Mega benchmark, nearly doubling the 11\.1x speedup of GLM\-5\.2\. This iteration also passed the single\-launch gate, a technical requirement that the previous version failed\. The performance gain relies on a CUDA \`load\_inline\` launch with persistent 512\-thread CTAs and 30 grid barriers\. The implementation keeps Int4 dequant inside the GEMV and absorbs Multi\-Head Latent Attention \(MLA\) so the fat KV is not materialized\.

## Sources

- [Story source](<https://huggingnews.com/ai/glm-53-hits-214x-pytorch-speedup-on-kernelbench-mega-doubling-glm-52-per-930bfc4c>)
- [Story source](<https://x.com/elliotarledge/status/2091125950243033209>)
- [Supporting source](<https://x.com/Zai_org/status/2091141282714587232>)

