MiMo-V2.6-Pro is Xiaomi MiMo’s flagship open-weight omnimodal model, built to scale reinforcement learning toward self-improvement across coding, agents, vision, and cybersecurity. Its sparse Mixture-of-Experts architecture contains 1.02T total parameters with 42B activated per token, using 384 routed experts with eight active per token. The model natively processes text, images, video, and audio and supports a 1M-token context window for large repositories, extended tool traces, and multi-session agent workflows. MiMo-V2.6-Pro-RL uses a unified mixed reinforcement learning process rather than separate domain-specific runs, alongside groupwise agentic grading that rewards higher-quality and more efficient solutions. Its architecture combines sliding-window and global attention, a 681M-parameter vision encoder, dedicated audio encoders, and a five-layer multi-token speculative decoder. It targets coding, general and visual agents, tool use, cybersecurity, long-horizon reasoning, etc.
Features
- 1.02T total parameters with 42B activated per token
- 384 routed experts with eight activated per token
- Native text, image, video, and audio understanding
- 1M-token context window for long-horizon workflows
- Unified reinforcement learning across multiple agent domains
- Groupwise agentic grading for self-improvement
- 681M vision encoder and dedicated audio encoders
- Five-layer multi-token speculative decoder for faster inference