---
format: "aidr-story-markdown/v1"
id: "a3487f668d2583376298fb2ae900e1da01332a4f42d12fd5479d8b463ffc4107"
canonical_url: "https://aidr.today/a3487f66?lang=en"
title: "DeepSeek Replaces Flagship V4-Pro With 552B Parameter V4.1-Flash Model"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-10T08:40:46.000Z"
category: "Models"
topics: ["deepseek","llm","multimodal","inference","benchmark","open-source","v41-flash","chips"]
source_urls: ["https://huggingnews.com/ai/deepseek-replaces-flagship-v4-pro-with-552b-parameter-v41-flash-model-eae1f63f","https://x.com/deepseek_ai/status/2097930608790167907","https://x.com/deepseek_ai/status/2097930613101838709","https://x.com/deepseek_ai/status/2097930617396887773","https://x.com/deepseek_ai/status/2097930620680941732","https://x.com/deepseek_ai/status/2097930624598446137","https://x.com/vllm_project/status/2097940813242405272","https://x.com/eliebakouch/status/2097931464344015347"]
summary: "A new Causal Encoder-Decoder architecture utilizes 8B active parameters for input and 16B for output in the latest AI release from DeepSeek. The 552B parameter Mixture-of-Experts model supports native multimodal input and a 1 million token context window. On September 14, 2026, all V4-Pro requests will route to the V4.1-Flash model at Flash rates, phasing out the Pro flagship after tests showed the new version leads in performance, cost, and speed. The system's KV cache requires 1/4 of the HBM and 1/8 of the SSD storage of its predecessor, reducing expenses for agent-based tasks. It incorporates a 197B parameter Engram lookup memory and was trained on 45 trillion tokens. Updated API pricing launched September 10, 2026, which maintains a 50% discount for workloads scheduled during off-peak hours."
---

# DeepSeek Replaces Flagship V4\-Pro With 552B Parameter V4\.1\-Flash Model

> [Open the canonical story](<https://aidr.today/a3487f66?lang=en>)

**Published:** 2026-09-10T08:40:46.000Z
**Category:** Models
**Topics:** deepseek, llm, multimodal, inference, benchmark, open\-source, v41\-flash, chips

## Summary

A new Causal Encoder\-Decoder architecture utilizes 8B active parameters for input and 16B for output in the latest AI release from DeepSeek\. The 552B parameter Mixture\-of\-Experts model supports native multimodal input and a 1 million token context window\. On September 14, 2026, all V4\-Pro requests will route to the V4\.1\-Flash model at Flash rates, phasing out the Pro flagship after tests showed the new version leads in performance, cost, and speed\. The system's KV cache requires 1/4 of the HBM and 1/8 of the SSD storage of its predecessor, reducing expenses for agent\-based tasks\. It incorporates a 197B parameter Engram lookup memory and was trained on 45 trillion tokens\. Updated API pricing launched September 10, 2026, which maintains a 50% discount for workloads scheduled during off\-peak hours\.

## Sources

- [Story source](<https://huggingnews.com/ai/deepseek-replaces-flagship-v4-pro-with-552b-parameter-v41-flash-model-eae1f63f>)
- [Story source](<https://x.com/deepseek_ai/status/2097930608790167907>)
- [Story source](<https://x.com/deepseek_ai/status/2097930613101838709>)
- [Story source](<https://x.com/deepseek_ai/status/2097930617396887773>)
- [Story source](<https://x.com/deepseek_ai/status/2097930620680941732>)
- [Story source](<https://x.com/deepseek_ai/status/2097930624598446137>)
- [Supporting source](<https://x.com/vllm_project/status/2097940813242405272>)
- [Supporting source](<https://x.com/eliebakouch/status/2097931464344015347>)

