Qwen3 is the large language model series developed by Qwen team
MiMo-V2-Flash: Efficient Reasoning, Coding, and Agentic Foundation
Moonshot's most powerful AI model
Qwen3-VL, the multimodal large language model series by Alibaba Cloud
Qwen3-Coder is the code version of Qwen3
A Next-Generation Training Engine Built for Ultra-Large MoE Models
A bidirectional pipeline parallelism algorithm
GLM-4.5: Open-source LLM for intelligent agents by Z.ai
gpt-oss-120b and gpt-oss-20b are two open-weight language models
Deep learning optimization library: makes distributed training easy
Open-source, high-performance Mixture-of-Experts large language model
Towards Ultimate Expert Specialization in Mixture-of-Experts Language
Building Mixture-of-Experts from LLaMA with Continual Pre-training
Large language model developed and released by NVIDIA
Open reasoning model for agentic coding and tool workflows
770B MoE model for coding, research, reasoning, and long-context work
Lightweight MoE model for local reasoning, coding, and AI agents
Efficient 30B MoE model for long-running agents and local inference
Efficient MoE model for reasoning, coding, and AI agent workflows
OpenAI’s compact 20B open model for fast, agentic, and local use
Efficient multimodal MoE model for coding, reasoning, and AI agents
Efficient 250B MoE model for agents, coding, and long-context work