Redundancy-aware KV Cache Compression for Reasoning Models
Unified KV Cache Compression Methods for Auto-Regressive Models
Mooncake is the serving platform for Kimi
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
FlashMLA: Efficient Multi-head Latent Attention Kernels
An expressive, efficient attention architecture
Claude + Obsidian knowledge companion
An open-source toolkit for BigMac-style pipeline-parallel training
Calculate token/s & GPU memory requirement for any LLM
Open agentic coding model optimized for local deployment
High-performance MoE model with MLA, MTP, and multilingual reasoning