---
format: "aidr-story-markdown/v1"
id: "f78fea501901d88dab01acb86726709eba621b2ca119af402babd8a185a8705a"
canonical_url: "https://aidr.today/f78fea50?lang=en"
title: "Alibaba Qwen Open Sources 6B Active Model as First Preview of Qwen4 Architecture"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-27T11:43:30.000Z"
category: "Releases"
topics: ["alibaba","qwen","open-source","llm","coding","inference"]
source_urls: ["https://huggingnews.com/ai/alibaba-qwen-open-sources-6b-active-model-as-first-preview-of-qwen4-arch-1add460b","https://x.com/Alibaba_Qwen/status/2092857839299826095","https://x.com/Alibaba_Qwen/status/2092867093385613760","https://x.com/ZhihuFrontier/status/2092865483989250537","https://x.com/QuixiAI/status/2092865018236711118"]
summary: "The company released its multimodal Qwen3.8-Flash model on OpenRouter and Qwen Cloud alongside the launch of Qwen3.8-Flash-Next. The open-weight Flash-Next model uses a mixture-of-experts design with 125B parameters, though only 6B are active per token. It incorporates 51B N-gram side parameters and utilizes a hybrid attention mechanism and the Muon optimizer to reduce compute costs. Qwen reported that Flash-Next outperforms DeepSeek-V4-Flash on coding and agent benchmarks despite having roughly half the active parameters. The model's API on Qwen Cloud is priced at $0.15 per 1M input tokens and $0.47 per 1M output tokens. This early release allows inference frameworks such as vLLM and SGLang to adapt to the new architecture before the formal Qwen4 launch."
---

# Alibaba Qwen Open Sources 6B Active Model as First Preview of Qwen4 Architecture

> [Open the canonical story](<https://aidr.today/f78fea50?lang=en>)

**Published:** 2026-08-27T11:43:30.000Z
**Category:** Releases
**Topics:** alibaba, qwen, open\-source, llm, coding, inference

## Summary

The company released its multimodal Qwen3\.8\-Flash model on OpenRouter and Qwen Cloud alongside the launch of Qwen3\.8\-Flash\-Next\. The open\-weight Flash\-Next model uses a mixture\-of\-experts design with 125B parameters, though only 6B are active per token\. It incorporates 51B N\-gram side parameters and utilizes a hybrid attention mechanism and the Muon optimizer to reduce compute costs\. Qwen reported that Flash\-Next outperforms DeepSeek\-V4\-Flash on coding and agent benchmarks despite having roughly half the active parameters\. The model's API on Qwen Cloud is priced at $0\.15 per 1M input tokens and $0\.47 per 1M output tokens\. This early release allows inference frameworks such as vLLM and SGLang to adapt to the new architecture before the formal Qwen4 launch\.

## Sources

- [Story source](<https://huggingnews.com/ai/alibaba-qwen-open-sources-6b-active-model-as-first-preview-of-qwen4-arch-1add460b>)
- [Story source](<https://x.com/Alibaba_Qwen/status/2092857839299826095>)
- [Story source](<https://x.com/Alibaba_Qwen/status/2092867093385613760>)
- [Supporting source](<https://x.com/ZhihuFrontier/status/2092865483989250537>)
- [Supporting source](<https://x.com/QuixiAI/status/2092865018236711118>)

