---
format: "aidr-story-markdown/v1"
id: "b07eb81cc5ac457ec819132baada751df64aa62e3d6f08ef93823af382604a45"
canonical_url: "https://aidr.today/b07eb81c?lang=en"
title: "Alibaba Releases Qwen3.8-Flash, First Open Weight Preview of Qwen4 Architecture"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-01T06:14:53.000Z"
category: "Models"
topics: ["alibaba","qwen","open-source","llm","agent","moe"]
source_urls: ["https://huggingnews.com/ai/alibaba-releases-qwen38-flash-first-open-weight-preview-of-qwen4-archite-1046cd12","https://x.com/arena/status/2094566204488962483","https://x.com/thenewstack/status/2094455234227585249"]
summary: "The production version of Alibaba's new multimodal model will cost $0.16 per 1 million input tokens and $0.47 per 1 million output tokens upon its release via QwenCloud API. The mixture-of-experts (MoE) design employs 125B parameters with 6B activated per token, using a hybrid of GDN and QSA attention to lower training expenses to 1/9 the cost of Qwen3.7-Plus. It supports a 262K native context window, which can be extended to 1 million tokens using YaRN. Qwen3.8-Flash-Next, an open-weight version of the new architecture, ranks No. 7 among open models and No. 24 overall in the Agent Arena based on 8,700 real-world agentic sessions. The model posted a 2.4% net improvement, surpassing the 1.5% improvement of Qwen3.8-27B. It specifically leads in task completion with a 12.3% confirmed success rate, ranking No. 5 among open models and No. 7 overall."
---

# Alibaba Releases Qwen3\.8\-Flash, First Open Weight Preview of Qwen4 Architecture

> [Open the canonical story](<https://aidr.today/b07eb81c?lang=en>)

**Published:** 2026-09-01T06:14:53.000Z
**Category:** Models
**Topics:** alibaba, qwen, open\-source, llm, agent, moe

## Summary

The production version of Alibaba's new multimodal model will cost $0\.16 per 1 million input tokens and $0\.47 per 1 million output tokens upon its release via QwenCloud API\. The mixture\-of\-experts \(MoE\) design employs 125B parameters with 6B activated per token, using a hybrid of GDN and QSA attention to lower training expenses to 1/9 the cost of Qwen3\.7\-Plus\. It supports a 262K native context window, which can be extended to 1 million tokens using YaRN\. Qwen3\.8\-Flash\-Next, an open\-weight version of the new architecture, ranks No\. 7 among open models and No\. 24 overall in the Agent Arena based on 8,700 real\-world agentic sessions\. The model posted a 2\.4% net improvement, surpassing the 1\.5% improvement of Qwen3\.8\-27B\. It specifically leads in task completion with a 12\.3% confirmed success rate, ranking No\. 5 among open models and No\. 7 overall\.

## Sources

- [Story source](<https://huggingnews.com/ai/alibaba-releases-qwen38-flash-first-open-weight-preview-of-qwen4-archite-1046cd12>)
- [Story source](<https://x.com/arena/status/2094566204488962483>)
- [Supporting source](<https://x.com/thenewstack/status/2094455234227585249>)

