DeepSeek-native AI coding agent for your terminal
Supercharge Your LLM with the Fastest KV Cache Layer
Redundancy-aware KV Cache Compression for Reasoning Models
Unified KV Cache Compression Methods for Auto-Regressive Models
Cache-Augmented Generation: A Simple, Efficient Alternative to RAG
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU
High-performance Inference and Deployment Toolkit for LLMs and VLMs
Mooncake is the serving platform for Kimi
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
Codex Switch & Instruct desktop manager
Trainable latent-memory framework for 100M-token contexts
Java wrapper for the popular chat & VOIP service
An expressive, efficient attention architecture
DeepSeek 4 Flash local inference engine for Metal
LLM inference server with continuous batching & SSD caching
Open-source TypeScript terminal coding agent for DeepSeek-V4
UCCL is an efficient communication library for GPUs
FlashMLA: Efficient Multi-head Latent Attention Kernels
An open-source AI agent that brings the power of Grok
A Model Context Protocol (MCP) Gateway & Registry
Claude Code, but it runs on your Mac for free
Claude + Obsidian knowledge companion
Chat with LLM like Vicuna totally in your browser with WebGPU
A course of learning LLM inference serving on Apple Silicon
The secure, validated skill registry for professional AI coding agents