MiMo-V2.6-Pro
1T omnimodal MoE model for coding, agents, and long-horizon reasoning
...Its sparse Mixture-of-Experts architecture contains 1.02T total parameters with 42B activated per token, using 384 routed experts with eight active per token. The model natively processes text, images, video, and audio and supports a 1M-token context window for large repositories, extended tool traces, and multi-session agent workflows. MiMo-V2.6-Pro-RL uses a unified mixed reinforcement learning process rather than separate domain-specific runs, alongside groupwise agentic grading that rewards higher-quality and more efficient solutions. Its architecture combines sliding-window and global attention, a 681M-parameter vision encoder, dedicated audio encoders, and a five-layer multi-token speculative decoder. ...