A library for accelerating Transformer models on NVIDIA GPUs
A real time inference engine for temporal logical specifications
LM Studio Apple MLX engine
High-performance reactive message-passing based Bayesian engine
DeepSeek 4 Flash local inference engine for Metal
A high-performance inference engine for AI models
A high-throughput and memory-efficient inference and serving engine
A lightweight vLLM implementation built from scratch
950 line, minimal, extensible LLM inference engine built from scratch
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU
LLM inference in C/C++
Jlama is a modern LLM inference engine for Java
Fastest LLM inference runtime for Apple Silicon
TokenSpeed is a speed-of-light LLM inference engine
Alibaba's high-performance LLM inference engine for diverse apps
High-performance inference framework for large language models
lightweight, standalone C++ inference engine for Google's Gemma models
Run the full 2.78-trillion-parameter Kimi K3 model
RGBD video generation model conditioned on camera input
Mooncake is the serving platform for Kimi
Fast Multimodal LLM on Mobile Devices
Code for running inference and finetuning with SAM 3 model
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine
Fast, flexible LLM inference
QVAC Fabric: cross-platform LLM inference and fine-tuning