Z AI released GLM-5.3 as a controlled experiment to demonstrate that scaling post-training and reinforcement learning (RL) can increase intelligence without increasing model size. The new model utilizes the same base, architecture, total parameters, and activated parameters as GLM-5.2, but underwent one month of scaling in long-horizon environments and RL. The company reported that the resulting performance gains were not marginal, identifying the post-training phase as the most effective lever for capability growth in the current development cycle.

The approach marks a reversal of the industry's previous trend of massive parameter growth, which Z AI described as a trillion-parameter detour. Historical analysis of the 1.6T parameter Switch Transformer, which used fewer than 3B activated parameters, showed that while large models excel at knowledge retrieval, they often fail at reasoning tasks. Z AI maintains that high-level inference, such as identifying cybersecurity vulnerabilities, requires scaling effective depth and post-training rather than increasing the volume of memorized facts.

Sign in to suggest edits

Key sources

  1. SOURCE@liamfedus“The 1.6T model achieved better C4 perplexities than the T5 models using far less compute, set a new SOTA on TriviaQA, but was dumb as bricks on reasoning tasks like SuperGLUE.”x.com
  2. SUPPORT@nandodf“This is the wisest and most accurate take on scaling laws — their promise and their costly failures — that you could read today.”x.com
  3. SOURCE@valsai“GLM 5.3 is #2 on Terminal Bench, #3 on Legal Bench, and #6 on Skills Bench among-open weight models”x.com
  4. SUPPORT@valsai“At just $0.31/test, it is $0.12 cheaper than 5.2”x.com
  5. SUPPORT@valsai“It is strongest on conclusion and issue-spotting tasks, with interpretation being the weak spot”x.com
  6. SUPPORT@valsai“We look forward to running our private benchmarks such as the Vals Index, Finance Agent, and Vibe Code Bench once the weights become available”x.com
Markdown