Run Local LLMs on Any Device. Open-source
FlashInfer: Kernel Library for LLM Serving
A high-throughput and memory-efficient inference and serving engine
Efficient few-shot learning with Sentence Transformers
AIMET is a library that provides advanced quantization and compression
A library for accelerating Transformer models on NVIDIA GPUs
Official inference library for Mistral models
Neural Network Compression Framework for enhanced OpenVINO
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
Training and deploying machine learning models on Amazon SageMaker
Uplift modeling and causal inference with machine learning algorithms
Large Language Model Text Generation Inference
Operating LLMs in production
Deep learning optimization library: makes distributed training easy
Everything you need to build state-of-the-art foundation models
Optimizing inference proxy for LLMs
Standardized Serverless ML Inference Platform on Kubernetes
The official Python client for the Huggingface Hub
LLM training code for MosaicML foundation models
Simplifies the local serving of AI models from any source
MII makes low-latency and high-throughput inference possible
Unified Model Serving Framework
Single-cell analysis in Python
DoWhy is a Python library for causal inference
An MLOps framework to package, deploy, monitor and manage models