Official inference library for Mistral models
Replace OpenAI GPT with another LLM in your app
High-performance inference server for text embeddings models API layer
Large Language Model Text Generation Inference
The Triton Inference Server provides an optimized cloud
Library for serving Transformers models on Amazon SageMaker
Standardized Serverless ML Inference Platform on Kubernetes
C++ library for high performance inference on NVIDIA GPUs
A high-throughput and memory-efficient inference and serving engine
Port of Facebook's LLaMA model in C/C++
Port of OpenAI's Whisper model in C/C++
LLM inference in C/C++
AirLLM 70B inference with single 4GB GPU
AlphaFold 3 inference pipeline
Deep learning optimization library: makes distributed training easy
A general-purpose probabilistic programming system
Low-latency REST API for serving text-embeddings
ONNX Runtime: cross-platform, high performance ML inferencing
High-performance reactive message-passing based Bayesian engine
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
Optimizing inference proxy for LLMs
Faster Whisper transcription with CTranslate2
FlashInfer: Kernel Library for LLM Serving
C#/.NET binding of llama.cpp, including LLaMa/GPT model inference
lightweight, standalone C++ inference engine for Google's Gemma models