Training and deploying machine learning models on Amazon SageMaker
Run Local LLMs on Any Device. Open-source
A high-throughput and memory-efficient inference and serving engine
FlashInfer: Kernel Library for LLM Serving
Single-cell analysis in Python
Ready-to-use OCR with 80+ supported languages
Python Package for ML-Based Heterogeneous Treatment Effects Estimation
DoWhy is a Python library for causal inference
Uplift modeling and causal inference with machine learning algorithms
The official Python client for the Huggingface Hub
Operating LLMs in production
A Pythonic framework to simplify AI service building
Everything you need to build state-of-the-art foundation models
Integrate, train and manage any AI models and APIs with your database
Neural Network Compression Framework for enhanced OpenVINO
Large Language Model Text Generation Inference
Official inference library for Mistral models
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
Optimizing inference proxy for LLMs
LMDeploy is a toolkit for compressing, deploying, and serving LLMs
MII makes low-latency and high-throughput inference possible
A unified framework for scalable computing
Sparsity-aware deep learning inference runtime for CPUs
Efficient few-shot learning with Sentence Transformers
A library for accelerating Transformer models on NVIDIA GPUs