---
format: "aidr-story-markdown/v1"
id: "b177ada21b6450258c3757f50d9d83a1bbfcd1013d3ba52fd6f29998b2f82635"
canonical_url: "https://aidr.today/b177ada2?lang=en"
title: "Z AI Scales GLM-5.3 Intelligence via RL, Reversing Trillion Parameter Detour"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-20T11:48:24.000Z"
category: "Models"
topics: ["llm","reasoning","fine-tuning","rl"]
source_urls: ["https://huggingnews.com/ai/z-ai-scales-glm-53-intelligence-via-rl-reversing-trillion-parameter-deto-dd880b3a","https://x.com/LiamFedus/status/2090363702042304847","https://x.com/NandoDF/status/2090333885858988299","https://x.com/ValsAI/status/2090192848780136668","https://x.com/ValsAI/status/2090192851057696924","https://x.com/ValsAI/status/2090192853033275460","https://x.com/ValsAI/status/2090192857902841931"]
summary: "Z AI released GLM-5.3 as a controlled experiment to demonstrate that scaling post-training and reinforcement learning (RL) can increase intelligence without increasing model size. The new model utilizes the same base, architecture, total parameters, and activated parameters as GLM-5.2, but underwent one month of scaling in long-horizon environments and RL. The company reported that the resulting performance gains were not marginal, identifying the post-training phase as the most effective lever for capability growth in the current development cycle. The approach marks a reversal of the industry's previous trend of massive parameter growth, which Z AI described as a trillion-parameter detour. Historical analysis of the 1.6T parameter Switch Transformer, which used fewer than 3B activated parameters, showed that while large models excel at knowledge retrieval, they often fail at reasoning tasks. Z AI maintains that high-level inference, such as identifying cybersecurity vulnerabilities, requires scaling effective depth and post-training rather than increasing the volume of memorized facts."
---

# Z AI Scales GLM\-5\.3 Intelligence via RL, Reversing Trillion Parameter Detour

> [Open the canonical story](<https://aidr.today/b177ada2?lang=en>)

**Published:** 2026-08-20T11:48:24.000Z
**Category:** Models
**Topics:** llm, reasoning, fine\-tuning, rl

## Summary

Z AI released GLM\-5\.3 as a controlled experiment to demonstrate that scaling post\-training and reinforcement learning \(RL\) can increase intelligence without increasing model size\. The new model utilizes the same base, architecture, total parameters, and activated parameters as GLM\-5\.2, but underwent one month of scaling in long\-horizon environments and RL\. The company reported that the resulting performance gains were not marginal, identifying the post\-training phase as the most effective lever for capability growth in the current development cycle\. The approach marks a reversal of the industry's previous trend of massive parameter growth, which Z AI described as a trillion\-parameter detour\. Historical analysis of the 1\.6T parameter Switch Transformer, which used fewer than 3B activated parameters, showed that while large models excel at knowledge retrieval, they often fail at reasoning tasks\. Z AI maintains that high\-level inference, such as identifying cybersecurity vulnerabilities, requires scaling effective depth and post\-training rather than increasing the volume of memorized facts\.

## Sources

- [Story source](<https://huggingnews.com/ai/z-ai-scales-glm-53-intelligence-via-rl-reversing-trillion-parameter-deto-dd880b3a>)
- [Story source](<https://x.com/LiamFedus/status/2090363702042304847>)
- [Supporting source](<https://x.com/NandoDF/status/2090333885858988299>)
- [Story source](<https://x.com/ValsAI/status/2090192848780136668>)
- [Supporting source](<https://x.com/ValsAI/status/2090192851057696924>)
- [Supporting source](<https://x.com/ValsAI/status/2090192853033275460>)
- [Supporting source](<https://x.com/ValsAI/status/2090192857902841931>)

