Colibrì is an open-source inference engine designed to run frontier-scale mixture-of-experts (MoE) models on consumer and heterogeneous hardware. It treats VRAM, system RAM, and NVMe storage as a unified memory hierarchy, dynamically placing and streaming model weights where they can be accessed most efficiently. This architecture enables users to run models ranging from 7 billion to 2.8 trillion parameters without requiring the entire model to fit in expensive GPU memory. Written primarily in pure C, Colibrì has no core engine dependencies and can operate without a GPU, although CUDA, Metal, Vulkan, and other acceleration options can improve performance. A consistent command-line interface, OpenAI-compatible API, web dashboard, and desktop application support multiple model families, including GLM, DeepSeek, Kimi, Qwen, Inkling, and OLMoE. Colibrì also serves as an open research platform for testing model placement, caching, compression, speculative decoding, storage I/O, & more.
Features
- Multi-tier memory management: Uses VRAM, RAM, and NVMe storage as a unified hierarchy for placing and streaming model weights.
- Frontier-scale model support: Runs MoE models ranging from compact 7B models to systems with as many as 2.8 trillion parameters.
- Broad hardware compatibility: Supports CPU-only operation as well as optional CUDA, Metal, Vulkan, HIP, NUMA, and multi-GPU acceleration.
- Intelligent expert caching: Applies LRU caching, learned hot-expert pinning, routing-based prefetching, and persistent usage history to improve performance.
- Unified interfaces: Provides terminal chat, an OpenAI-compatible API, a browser-based dashboard, a desktop application, and local cluster capabilities.
- Correctness-focused optimization: Preserves model precision and routing semantics while measuring throughput, latency, memory use, quality, and hardware cost.