A library for accelerating Transformer models on NVIDIA GPUs
A real time inference engine for temporal logical specifications
LM Studio Apple MLX engine
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU
Run the full 2.78-trillion-parameter Kimi K3 model
Run frontier MoE models on hardware you already own
DeepEP: an efficient expert-parallel communication library
High-performance reactive message-passing based Bayesian engine
Ling is a MoE LLM provided and open-sourced by InclusionAI
DeepSeek 4 Flash local inference engine for Metal
A high-throughput and memory-efficient inference and serving engine
Open-source large language model family from Tencent Hunyuan
A lightweight vLLM implementation built from scratch
Ring is a reasoning MoE LLM provided and open-sourced by InclusionAI
950 line, minimal, extensible LLM inference engine built from scratch
A high-performance inference engine for AI models
Deep learning optimization library: makes distributed training easy
lightweight, standalone C++ inference engine for Google's Gemma models
Jlama is a modern LLM inference engine for Java
Blazing fast, instant realtime GraphQL APIs on your DB
TokenSpeed is a speed-of-light LLM inference engine
High-performance inference framework for large language models
QVAC Fabric: cross-platform LLM inference and fine-tuning
LLM inference in C/C++
Alibaba's high-performance LLM inference engine for diverse apps