The company released its multimodal Qwen3.8-Flash model on OpenRouter and Qwen Cloud alongside the launch of Qwen3.8-Flash-Next. The open-weight Flash-Next model uses a mixture-of-experts design with 125B parameters, though only 6B are active per token. It incorporates 51B N-gram side parameters and utilizes a hybrid attention mechanism and the Muon optimizer to reduce compute costs.

Qwen reported that Flash-Next outperforms DeepSeek-V4-Flash on coding and agent benchmarks despite having roughly half the active parameters. The model's API on Qwen Cloud is priced at $0.15 per 1M input tokens and $0.47 per 1M output tokens. This early release allows inference frameworks such as vLLM and SGLang to adapt to the new architecture before the formal Qwen4 launch.

Sign in to suggest edits

Key sources

  1. SOURCE@alibaba_qwen“Coding assistants, agentic workflows, long-video understanding, all one API call away”x.com
  2. SOURCE@alibaba_qwen“$0.15/1M input tokens, $0.47/1M output tokens, and just $0.016/1M on cache hits”x.com
  3. SUPPORT@zhihufrontier“Flash-Next has 125B total parameters with only 6B active — far smaller than the 300B-class GLM-5.3-Flash”x.com
  4. SUPPORT@quixiai“running at 150 tok/s @ c1 to 661.1 tok/s @ c32 at 262k context”x.com
Markdown