Colibri is a compact inference engine designed to run the 744-billion-parameter GLM-5.2 mixture-of-experts model on consumer hardware. It keeps the dense portion of the quantized model in memory while streaming routed experts from a large disk-based store as they are needed. The runtime is implemented in pure C, requires no Python or BLAS during inference, and can operate without a GPU. Compressed attention caches, expert caching, optional hot tiers, and speculative decoding reduce memory pressure and improve repeated use. A planning tool calculates safe disk, RAM, and VRAM placement before loading the model, while a diagnostic command checks system readiness. Colibri includes terminal chat, an OpenAI-compatible text API, and a browser client, but disk-bound generation can be slow on cold caches.

Features

  • GLM-5.2 inference on consumer hardware
  • Disk-streamed mixture-of-experts architecture
  • Dependency-free pure C inference runtime
  • Compressed KV cache and expert caching
  • Automatic RAM and VRAM placement planning
  • Terminal chat and OpenAI-compatible API

Project Samples

Project Activity

See All Activity >

Categories

AI Models

License

Apache License V2.0

Follow Colibrì

Colibrì Web Site

Other Useful Business Software
Build Agents and Models on One Platform Icon
Build Agents and Models on One Platform

Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
Try It Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of Colibrì!

Additional Project Details

Operating Systems

Linux, Mac, Windows

Programming Language

C

Related Categories

C AI Models

Registered

2026-07-13