Official inference library for Mistral models
Replace OpenAI GPT with another LLM in your app
High-performance inference server for text embeddings models API layer
Large Language Model Text Generation Inference
A high-throughput and memory-efficient inference and serving engine
Library for serving Transformers models on Amazon SageMaker
AirLLM 70B inference with single 4GB GPU
Optimizing inference proxy for LLMs
QVAC Fabric: cross-platform LLM inference and fine-tuning
Deep learning optimization library: makes distributed training easy
Port of Facebook's LLaMA model in C/C++
Port of OpenAI's Whisper model in C/C++
High-performance Inference and Deployment Toolkit for LLMs and VLMs
LLM inference in C/C++
A general-purpose probabilistic programming system
Python-free Rust inference server
C#/.NET binding of llama.cpp, including LLaMa/GPT model inference
ONNX Runtime: cross-platform, high performance ML inferencing
High-performance reactive message-passing based Bayesian engine
Bayesian inference with probabilistic programming
LightLLM is a Python-based LLM (Large Language Model) inference
950 line, minimal, extensible LLM inference engine built from scratch
DeepSeek 4 Flash local inference engine for Metal
LLM inference server with continuous batching & SSD caching
lightweight, standalone C++ inference engine for Google's Gemma models