Qwen3.8-2.4T-A95B is Qwen’s largest open-weight model and the first Qwen-Max-class model released openly, targeting advanced coding, professional work, research, and long-horizon agentic tasks. It uses a massive Mixture-of-Experts architecture with 2.4 trillion total parameters while activating 95B per token, combining Gated DeltaNet and attention layers across 512 experts. The model emphasizes reliable autonomous execution, including stronger planning, environment feedback handling, and end-to-end completion of complex multi-step workflows. It natively supports a 262K-token context window that can be extended beyond one million tokens. Qwen3.8 also provides adjustable reasoning depth through low, medium, and xhigh reasoning-effort settings and preserves reasoning context across conversations. It is a text-only, thinking-first model and supports deployment through vLLM, SGLang, and TokenSpeed, with strong results across coding, tool use, research, and professional benchmarks.
Features
- 2.4T total parameters with 95B activated per token
- 512-expert Mixture-of-Experts architecture
- Hybrid Gated DeltaNet and attention architecture
- 262K native context, extensible beyond 1M tokens
- Advanced coding and repository-level software engineering
- Autonomous planning and long-horizon agent execution
- Adjustable low, medium, and xhigh reasoning effort
- Preserve-thinking support for multi-turn reasoning context