The company released its multimodal Qwen3.8-Flash model on OpenRouter and Qwen Cloud alongside the launch of Qwen3.8-Flash-Next. The open-weight Flash-Next model uses a mixture-of-experts design with 125B parameters, though only 6B are active per token. It incorporates 51B N-gram side parameters and utilizes a hybrid attention mechanism and the Muon optimizer to reduce compute costs.
Qwen reported that Flash-Next outperforms DeepSeek-V4-Flash on coding and agent benchmarks despite having roughly half the active parameters. The model's API on Qwen Cloud is priced at $0.15 per 1M input tokens and $0.47 per 1M output tokens. This early release allows inference frameworks such as vLLM and SGLang to adapt to the new architecture before the formal Qwen4 launch.
Key sources
- SOURCE@alibaba_qwen“Coding assistants, agentic workflows, long-video understanding, all one API call away”x.com
- SOURCE@alibaba_qwen“$0.15/1M input tokens, $0.47/1M output tokens, and just $0.016/1M on cache hits”x.com
- SUPPORT@zhihufrontier“Flash-Next has 125B total parameters with only 6B active — far smaller than the 300B-class GLM-5.3-Flash”x.com
- SUPPORT@quixiai“running at 150 tok/s @ c1 to 661.1 tok/s @ c32 at 262k context”x.com