Redundancy-aware KV Cache Compression for Reasoning Models
Unified KV Cache Compression Methods for Auto-Regressive Models
Supercharge Your LLM with the Fastest KV Cache Layer
Claude + Obsidian knowledge companion
LLM inference server with continuous batching & SSD caching
Claude Code, but it runs on your Mac for free
TensorRT LLM provides users with an easy-to-use Python API
An open-source toolkit for BigMac-style pipeline-parallel training
Advancing Open-source World Models
RNN with great LLM performance