Port of OpenAI's Whisper model in C/C++
Bring the notion of Model-as-a-Service to life
AIMET is a library that provides advanced quantization and compression
Library for serving Transformers models on Amazon SageMaker
Uncover insights, surface problems, monitor, and fine tune your LLM
INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model
PyTorch extensions for fast R&D prototyping and Kaggle farming
Probabilistic reasoning and statistical analysis in TensorFlow
Official inference library for Mistral models
Deep learning optimization library: makes distributed training easy
High-performance neural network inference framework for mobile
The official Python client for the Huggingface Hub
Fast inference engine for Transformer models
A unified framework for scalable computing
State-of-the-art diffusion models for image and audio generation
Pytorch domain library for recommendation systems
A lightweight vision library for performing large object detection
C++ library for high performance inference on NVIDIA GPUs
Replace OpenAI GPT with another LLM in your app
A general-purpose probabilistic programming system
MII makes low-latency and high-throughput inference possible
Bolt is a deep learning library with high performance
PyTorch library of curated Transformer models and their components
An innovative library for efficient LLM inference
Lightweight inference library for ONNX files, written in C++