MiMo-V2.6-Flash is Xiaomi MiMo’s efficiency-balanced open-weight omnimodal model, designed to scale reinforcement learning across coding, general agents, visual tasks, and cybersecurity. Its sparse Mixture-of-Experts architecture contains 309B total parameters while activating only 15B per token, using 256 routed experts with eight active per token. The model natively processes text, images, video, and audio and supports a 1M-token context window for large repositories, extended tool traces, and multi-session agent workflows. Training uses a unified mixed RL process rather than separate domain-specific runs, alongside asynchronous GRPO and groupwise agentic grading that rewards higher-quality and more efficient solutions. Its architecture combines sliding-window and global attention with a 681M-parameter vision encoder and dedicated audio encoders. A five-layer speculative decoder predicts multiple subsequent tokens to accelerate inference.
Features
- 309B total parameters with 15B activated per token
- 256 routed experts with eight activated per token
- Native text, image, video, and audio understanding
- 1M-token context window for long-horizon workflows
- Unified reinforcement learning across multiple agent domains
- Groupwise agentic grading for self-improvement
- 681M vision encoder and dedicated audio encoders
- Five-layer multi-token speculative decoder for faster inference