MiMo-V2.5 is a native omnimodal large language model developed by Xiaomi, designed for advanced agentic workflows, multimodal reasoning, and long-context processing. Built on a Mixture-of-Experts architecture with approximately 309B total parameters and around 15B activated per inference, it balances high capability with efficient execution. The model natively processes text, images, video, and audio within a unified system, enabling cross-modal understanding and complex task execution in a single pipeline. With a context window of up to 1 million tokens, it can handle large documents, extended conversations, and multi-step workflows without fragmentation. MiMo-V2.5 delivers near-Pro-level performance in coding, reasoning, and agent tasks while maintaining lower cost and faster inference speeds. It also integrates advanced components such as multi-token prediction modules and specialized vision and audio encoders, making it well-suited for autonomous agents and software development.
Features
- Native omnimodal support for text, image, video, and audio
- Mixture-of-Experts architecture with ~309B total parameters
- ~15B active parameters for efficient inference
- 1M-token context window for long-horizon tasks
- Strong agentic performance close to Pro-level models
- Multi-token prediction for faster generation
- Integrated vision transformer and audio encoder
- Optimized for autonomous workflows and tool-based execution