MiMo-V2.6-Pro
1T omnimodal MoE model for coding, agents, and long-horizon reasoning
...The model natively processes text, images, video, and audio and supports a 1M-token context window for large repositories, extended tool traces, and multi-session agent workflows. MiMo-V2.6-Pro-RL uses a unified mixed reinforcement learning process rather than separate domain-specific runs, alongside groupwise agentic grading that rewards higher-quality and more efficient solutions. Its architecture combines sliding-window and global attention, a 681M-parameter vision encoder, dedicated audio encoders, and a five-layer multi-token speculative decoder. It targets coding, general and visual agents, tool use, cybersecurity, long-horizon reasoning, etc.