The production version of Alibaba's new multimodal model will cost $0.16 per 1 million input tokens and $0.47 per 1 million output tokens upon its release via QwenCloud API. The mixture-of-experts (MoE) design employs 125B parameters with 6B activated per token, using a hybrid of GDN and QSA attention to lower training expenses to 1/9 the cost of Qwen3.7-Plus. It supports a 262K native context window, which can be extended to 1 million tokens using YaRN.
Qwen3.8-Flash-Next, an open-weight version of the new architecture, ranks No. 7 among open models and No. 24 overall in the Agent Arena based on 8,700 real-world agentic sessions. The model posted a 2.4% net improvement, surpassing the 1.5% improvement of Qwen3.8-27B. It specifically leads in task completion with a 12.3% confirmed success rate, ranking No. 5 among open models and No. 7 overall.