Redundancy-aware KV Cache Compression for Reasoning Models
Unified KV Cache Compression Methods for Auto-Regressive Models
Supercharge Your LLM with the Fastest KV Cache Layer
Claude + Obsidian knowledge companion
LLM inference server with continuous batching & SSD caching
a pluggable app that runs a full check on the deployment
Claude Code, but it runs on your Mac for free
TensorRT LLM provides users with an easy-to-use Python API
An open-source toolkit for BigMac-style pipeline-parallel training
Advancing Open-source World Models
Serverless Python
SweptPC — Free Portable Windows PC Cleaner
RNN with great LLM performance
Serverless Python
Version controlled file system
Numerical Transient Simulator for Power System
e500v2 simulator
Python Killboard Platform for EVE Online
High performance distributed in-memory key/value store
Find duplicate videos by content