FlashInfer: Kernel Library for LLM Serving
A high-throughput and memory-efficient inference and serving engine
Run Local LLMs on Any Device. Open-source
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
Efficient few-shot learning with Sentence Transformers
AIMET is a library that provides advanced quantization and compression
Sparsity-aware deep learning inference runtime for CPUs
The official Python client for the Huggingface Hub
Everything you need to build state-of-the-art foundation models
Uplift modeling and causal inference with machine learning algorithms
A library for accelerating Transformer models on NVIDIA GPUs
A set of Docker images for training and serving models in TensorFlow
Uncover insights, surface problems, monitor, and fine tune your LLM
Neural Network Compression Framework for enhanced OpenVINO
State-of-the-art diffusion models for image and audio generation
Official inference library for Mistral models
Scripts for fine-tuning Meta Llama3 with composable FSDP & PEFT method
Unified Model Serving Framework
LLM training code for MosaicML foundation models
Optimizing inference proxy for LLMs
A Pythonic framework to simplify AI service building
DoWhy is a Python library for causal inference
Training and deploying machine learning models on Amazon SageMaker
Replace OpenAI GPT with another LLM in your app
An MLOps framework to package, deploy, monitor and manage models