Zai's latest model utilized Kimi-Linear Decode on an RTX PRO 6000 GPU to optimize inference throughput. The GLM-5.3 version reached 21.4x the PyTorch baseline on the KernelBench-Mega benchmark, nearly doubling the 11.1x speedup of GLM-5.2. This iteration also passed the single-launch gate, a technical requirement that the previous version failed.

The performance gain relies on a CUDA `load_inline` launch with persistent 512-thread CTAs and 30 grid barriers. The implementation keeps Int4 dequant inside the GEMV and absorbs Multi-Head Latent Attention (MLA) so the fat KV is not materialized.

Sign in to suggest edits

Key sources

  1. SOURCE@elliotarledge“MLA is absorbed so the fat KV is not materialized”x.com
  2. SUPPORT@zai_org“One CUDA `load_inline` launch”x.com
Markdown