A library for accelerating Transformer models on NVIDIA GPUs
A real time inference engine for temporal logical specifications
LM Studio Apple MLX engine
High-performance reactive message-passing based Bayesian engine
DeepSeek 4 Flash local inference engine for Metal
A high-performance inference engine for AI models
A high-throughput and memory-efficient inference and serving engine
A lightweight vLLM implementation built from scratch
950 line, minimal, extensible LLM inference engine built from scratch
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU
LLM inference in C/C++
Jlama is a modern LLM inference engine for Java
Fastest LLM inference runtime for Apple Silicon
TokenSpeed is a speed-of-light LLM inference engine
Alibaba's high-performance LLM inference engine for diverse apps
High-performance inference framework for large language models
lightweight, standalone C++ inference engine for Google's Gemma models
Run the full 2.78-trillion-parameter Kimi K3 model
RGBD video generation model conditioned on camera input
Mooncake is the serving platform for Kimi
Low-latency AI inference engine optimized for mobile devices
Fast Multimodal LLM on Mobile Devices
Code for running inference and finetuning with SAM 3 model
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine
Fast, flexible LLM inference