Port of OpenAI's Whisper model in C/C++
Bring the notion of Model-as-a-Service to life
AIMET is a library that provides advanced quantization and compression
Library for serving Transformers models on Amazon SageMaker
Uncover insights, surface problems, monitor, and fine tune your LLM
Official inference library for Mistral models
INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model
PyTorch extensions for fast R&D prototyping and Kaggle farming
Deep learning optimization library: makes distributed training easy
Probabilistic reasoning and statistical analysis in TensorFlow
High-performance neural network inference framework for mobile
The official Python client for the Huggingface Hub
A unified framework for scalable computing
A lightweight vision library for performing large object detection
Fast inference engine for Transformer models
State-of-the-art diffusion models for image and audio generation
Pytorch domain library for recommendation systems
MII makes low-latency and high-throughput inference possible
Replace OpenAI GPT with another LLM in your app
A general-purpose probabilistic programming system
Bolt is a deep learning library with high performance
PyTorch library of curated Transformer models and their components
An innovative library for efficient LLM inference
Lightweight inference library for ONNX files, written in C++
Serve machine learning models within a Docker container