---
format: "aidr-story-markdown/v1"
id: "aa844bb57668df6146f63e7a397fac67126f02b010153ae652403b163d404e44"
canonical_url: "https://aidr.today/aa844bb5?lang=en"
title: "SGLang Cuts 1T Model Restarts to 32 Seconds, Down From 8.8 Minutes"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-22T08:48:35.000Z"
category: "Infra"
topics: ["alibaba","infra","inference","llm"]
source_urls: ["https://huggingnews.com/ai/sglang-cuts-1t-model-restarts-to-32-seconds-down-from-88-minutes-230119fa","https://x.com/AntLingAGI/status/2091021795373855124","https://x.com/sgl_project/status/2091048937558085862"]
summary: "AntLingAGI and Alibaba developed a new memory optimization tool to accelerate the deployment of massive language models. The Weight Cache Daemon for SGLang reduced weight loading for the Ling-2.6-1T FP8 model to 0.63 seconds, which is up to 780 times faster than loading from disk. This optimization cuts total engine startup to 0.53 minutes, enabling the restart of a trillion parameter model in 32 seconds. The tool targets the bottleneck of weight loading in production environments to improve recovery times and resource efficiency. SGLang serves as the serving framework for the deployment, and the daemon allows operators to bypass slow disk-based loading by caching model weights in memory."
---

# SGLang Cuts 1T Model Restarts to 32 Seconds, Down From 8\.8 Minutes

> [Open the canonical story](<https://aidr.today/aa844bb5?lang=en>)

**Published:** 2026-08-22T08:48:35.000Z
**Category:** Infra
**Topics:** alibaba, infra, inference, llm

## Summary

AntLingAGI and Alibaba developed a new memory optimization tool to accelerate the deployment of massive language models\. The Weight Cache Daemon for SGLang reduced weight loading for the Ling\-2\.6\-1T FP8 model to 0\.63 seconds, which is up to 780 times faster than loading from disk\. This optimization cuts total engine startup to 0\.53 minutes, enabling the restart of a trillion parameter model in 32 seconds\. The tool targets the bottleneck of weight loading in production environments to improve recovery times and resource efficiency\. SGLang serves as the serving framework for the deployment, and the daemon allows operators to bypass slow disk\-based loading by caching model weights in memory\.

## Sources

- [Story source](<https://huggingnews.com/ai/sglang-cuts-1t-model-restarts-to-32-seconds-down-from-88-minutes-230119fa>)
- [Story source](<https://x.com/AntLingAGI/status/2091021795373855124>)
- [Story source](<https://x.com/sgl_project/status/2091048937558085862>)

