Run Local LLMs on Any Device. Open-source
FlashInfer: Kernel Library for LLM Serving
A high-throughput and memory-efficient inference and serving engine
Ready-to-use OCR with 80+ supported languages
Operating LLMs in production
Neural Network Compression Framework for enhanced OpenVINO
Official inference library for Mistral models
Efficient few-shot learning with Sentence Transformers
An MLOps framework to package, deploy, monitor and manage models
LLM training code for MosaicML foundation models
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
Everything you need to build state-of-the-art foundation models
Optimizing inference proxy for LLMs
The official Python client for the Huggingface Hub
Training and deploying machine learning models on Amazon SageMaker
Large Language Model Text Generation Inference
A library for accelerating Transformer models on NVIDIA GPUs
LMDeploy is a toolkit for compressing, deploying, and serving LLMs
Simplifies the local serving of AI models from any source
MII makes low-latency and high-throughput inference possible
Library for OCR-related tasks powered by Deep Learning
AIMET is a library that provides advanced quantization and compression
Uplift modeling and causal inference with machine learning algorithms
Single-cell analysis in Python
Deep learning optimization library: makes distributed training easy