Official inference library for Mistral models
Replace OpenAI GPT with another LLM in your app
High-performance inference server for text embeddings models API layer
The Triton Inference Server provides an optimized cloud
Large Language Model Text Generation Inference
Library for serving Transformers models on Amazon SageMaker
A high-throughput and memory-efficient inference and serving engine
AirLLM 70B inference with single 4GB GPU
Optimizing inference proxy for LLMs
QVAC Fabric: cross-platform LLM inference and fine-tuning
Port of OpenAI's Whisper model in C/C++
Deep learning optimization library: makes distributed training easy
Port of Facebook's LLaMA model in C/C++
C++ library for high performance inference on NVIDIA GPUs
High-performance Inference and Deployment Toolkit for LLMs and VLMs
Python-free Rust inference server
A general-purpose probabilistic programming system
C#/.NET binding of llama.cpp, including LLaMa/GPT model inference
Faster Whisper transcription with CTranslate2
ONNX Runtime: cross-platform, high performance ML inferencing
LightLLM is a Python-based LLM (Large Language Model) inference
950 line, minimal, extensible LLM inference engine built from scratch
High-performance reactive message-passing based Bayesian engine
Bayesian inference with probabilistic programming
DeepSeek 4 Flash local inference engine for Metal