Official inference library for Mistral models
Replace OpenAI GPT with another LLM in your app
High-performance inference server for text embeddings models API layer
The Triton Inference Server provides an optimized cloud
Large Language Model Text Generation Inference
A high-throughput and memory-efficient inference and serving engine
Library for serving Transformers models on Amazon SageMaker
FlashInfer: Kernel Library for LLM Serving
AirLLM 70B inference with single 4GB GPU
Optimizing inference proxy for LLMs
QVAC Fabric: cross-platform LLM inference and fine-tuning
AlphaFold 3 inference pipeline
Deep learning optimization library: makes distributed training easy
C++ library for high performance inference on NVIDIA GPUs
Port of Facebook's LLaMA model in C/C++
Port of OpenAI's Whisper model in C/C++
High-performance Inference and Deployment Toolkit for LLMs and VLMs
LLM inference in C/C++
A general-purpose probabilistic programming system
Python-free Rust inference server
C#/.NET binding of llama.cpp, including LLaMa/GPT model inference
A high-performance inference system for large language models
Faster Whisper transcription with CTranslate2
ONNX Runtime: cross-platform, high performance ML inferencing
Bayesian inference with probabilistic programming