A new Causal Encoder-Decoder architecture utilizes 8B active parameters for input and 16B for output in the latest AI release from DeepSeek. The 552B parameter Mixture-of-Experts model supports native multimodal input and a 1 million token context window. On September 14, 2026, all V4-Pro requests will route to the V4.1-Flash model at Flash rates, phasing out the Pro flagship after tests showed the new version leads in performance, cost, and speed.
The system's KV cache requires 1/4 of the HBM and 1/8 of the SSD storage of its predecessor, reducing expenses for agent-based tasks. It incorporates a 197B parameter Engram lookup memory and was trained on 45 trillion tokens. Updated API pricing launched September 10, 2026, which maintains a 50% discount for workloads scheduled during off-peak hours.
Key sources
- SOURCE@deepseek_ai“smallest model in our new architecture family, with native visual understanding”x.com
- SOURCE@deepseek_ai“Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output”x.com
- SOURCE@deepseek_ai“KV cache needs just: 1/4 the HBM 1/8 the SSD storage”x.com
- SOURCE@deepseek_ai“all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates”x.com
- SOURCE@deepseek_ai“Off-peak rates are 50% of peak rates”x.com
- SUPPORT@vllm_project“Engram: a quarter of the checkpoint is n-gram memory the model looks up instead of computes”x.com
- SUPPORT@eliebakouch“trained on 45T tokens”x.com
- SUPPORT@reuters“China's DeepSeek launches V4.1-Flash model”x.com