← Back to live feed1 story
1
GLM-5.3 Hits 21.4x PyTorch Speedup on KernelBench-Mega, Doubling GLM-5.2 Performance
Research41d agoZai's latest model utilized Kimi-Linear Decode on an RTX PRO 6000 GPU to optimize inference throughput. The GLM-5.3 version reached 21.4x the PyTorch baseline on the KernelBench-Mega benchmark, nearly doubling the 11.1x speedup of GLM-5.2. This iteration also passed the single-launch gate, a technical requirement that the previous version failed.
The performance gain relies on a CUDA `load_inline` launch with persistent 512-thread CTAs and 30 grid barriers. The implementation keeps Int4 dequant inside the GEMV and absorbs Multi-Head Latent Attention (MLA) so the fat KV is not materialized.
Sign in to suggest edits