Kolibri-1 is Aleph Alpha’s open-weight bilingual Mixture-of-Experts reasoning model, designed for efficient German- and English-language AI systems. It contains 78B total parameters while activating only 3.46B per token, providing substantial model capacity with relatively low inference compute. Its 50-layer architecture uses 384 experts per layer, with six routed experts and one shared expert active during processing, alongside a 4:1 combination of sliding-window and grouped-query attention. Kolibri supports explicit reasoning with configurable low, medium, and high effort levels, as well as structured tool calling for agentic workflows. Its native 262K-token context can scale to a validated 1M tokens, enabling long-document processing, RAG, research, and extended coding tasks. The model was pretrained on 20T bilingual and code tokens, followed by additional mid-training and long-context training. FP8 weights reduce its memory footprint to approximately 78 GB.
Features
- 78B total parameters with 3.46B active per token
- 384 experts with six routed plus one shared expert active
- Native 262K context validated up to 1M tokens
- German and English bilingual specialization
- Low, medium, and high configurable reasoning effort
- Structured tool calling for agentic workflows
- Hybrid sliding-window and grouped-query attention
- FP8 weights with approximately 78 GB memory footprint