Qwen3.8-2.4T-A95B is Qwen’s largest open-weight model and the first Qwen-Max-class model released openly, targeting advanced coding, professional work, research, and long-horizon agentic tasks. It uses a massive Mixture-of-Experts architecture with 2.4 trillion total parameters while activating 95B per token, combining Gated DeltaNet and attention layers across 512 experts. The model emphasizes reliable autonomous execution, including stronger planning, environment feedback handling, and end-to-end completion of complex multi-step workflows. It natively supports a 262K-token context window that can be extended beyond one million tokens. ...