CompaniesOtherKey event

Alibaba Qwen Open Sources Qwen3.8-2.4T-A95B Flagship Model, 2.4 Trillion Parameters, 91% Size Reduction

Published: Updated: By 24TopNews Editorial Desk

Alibaba's Qwen team has released the Qwen3.8-2.4T-A95B model, featuring 2.4 trillion total parameters with 95 billion activated per token and native support for 262,144 token context. The model architecture is based on Qwen3.5 and demonstrates performance gains in programming, office tasks, research, and long-horizon agent applications. This is the first time Qwen has open-sourced a Max-level flagship model. The model size has been compressed by 91% from 4.9TB to 397GB using dynamic 1-bit hierarchical selective quantization technology from Unsloth AI. Benchmark results show leadership in PaperBench (93.0 points) and OSWorld-Verified (86.1 points), while lagging in TerminalBench 2.1 (86.6 points), SWE-bench Pro (67.7 points), and VideoMME v2 (68.3 points). API pricing is RMB 12 per million input tokens and RMB 36 per million output tokens domestically, with overseas pricing at USD 2 and USD 6 respectively.

Alibaba's Qwen team has open-sourced the weights of the Qwen3.8-2.4T-A95B model. The model has 2.4 trillion total parameters, with 95 billion parameters activated per token. It natively supports a context length of 262,144 tokens, expandable to 1.01 million tokens. Built on the Qwen3.5 architecture, the model delivers enhanced performance in programming, office productivity, scientific research, and long-cycle agent tasks. This marks the first time Qwen has made Max-level flagship model weights openly available.

Model deployment supports the SGLang, vLLM, and TokenSpeed inference engines, requiring the latest Recipe from each framework. The model operates by default in thinking mode, which cannot be disabled. The open-source AI project Unsloth AI applied dynamic 1-bit hierarchical selective quantization technology to compress the model size from 4.9TB to 397GB, achieving a 91% reduction in storage footprint. Using the Unsloth-Desktop tool, the model can run locally on devices with a combined memory and video memory of at least 410GB.

In benchmark tests, Qwen3.8-Max scored 93.0 points on PaperBench, surpassing GPT-5.6 Sol, Fable 5, and Claude Opus 4.8. It achieved the top score of 86.1 points on OSWorld-Verified and 91.5 points on the parameterized CAD benchmark, outperforming Fable 5 and others. However, it scored 86.6 points on TerminalBench 2.1, below GPT-5.6 Sol's 88.8 points; 67.7 points on SWE-bench Pro, trailing Fable 5's 80.0 points; and 68.3 points on VideoMME v2, below GPT-5.6 Sol's 71.1 points.

The Qwen team tested the model's autonomous task-completion capabilities, including continuous autonomous programming for approximately 16 days, independent reproduction of research papers, participation in algorithm competitions, and operation of a virtual e-commerce enterprise. During the autonomous programming task, the model built a self-evolving Harness to perform requirement collection, code generation, and bug fixing.

Regarding API pricing, Qwen3.8-Max charges RMB 12 per million input tokens and RMB 36 per million output tokens domestically, with implicit cache hits at RMB 1.5 per million tokens. Overseas pricing is USD 2 per million input tokens and USD 6 per million output tokens, with implicit cache hits at USD 0.25 per million tokens.

24TOPNEWS IMPACT INTELLIGENCE

Why this event matters

The event has a measured impact on 1 industry. The strongest current signal is positive for Artificial Intelligence, with intensity 80/100 and 90% confidence over a short term horizon.

Technology · 10.4

Artificial Intelligence

Direction
positive
Intensity
80
Confidence
90%
Horizon
Short term
Effective impact +61

Impact figures are analytical estimates that combine direction, intensity, confidence and event importance. They are not investment advice.