Redundancy-aware KV Cache Compression for Reasoning Models
Unified KV Cache Compression Methods for Auto-Regressive Models
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU
DeepSeek-native AI coding agent for your terminal
Mooncake is the serving platform for Kimi
Open-source TypeScript terminal coding agent for DeepSeek-V4
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
FlashMLA: Efficient Multi-head Latent Attention Kernels
An expressive, efficient attention architecture
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine
Claude + Obsidian knowledge companion
Chat, serve, monitor, and connect MLX models from one macOS app
Chat with LLM like Vicuna totally in your browser with WebGPU
LLM inference server with continuous batching & SSD caching
Run the full 2.78-trillion-parameter Kimi K3 model
Claude Code, but it runs on your Mac for free
TensorRT LLM provides users with an easy-to-use Python API
An open-source toolkit for BigMac-style pipeline-parallel training
Advancing Open-source World Models
RNN with great LLM performance
Calculate token/s & GPU memory requirement for any LLM
Open agentic coding model optimized for local deployment
High-performance MoE model with MLA, MTP, and multilingual reasoning