Run Local LLMs on Any Device. Open-source
Standardized Serverless ML Inference Platform on Kubernetes
A high-throughput and memory-efficient inference and serving engine
Low-latency REST API for serving text-embeddings
Trainable models and NN optimization tools
Everything you need to build state-of-the-art foundation models
FlashInfer: Kernel Library for LLM Serving
State-of-the-art diffusion models for image and audio generation
Sparsity-aware deep learning inference runtime for CPUs
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
AIMET is a library that provides advanced quantization and compression
Libraries for applying sparsification recipes to neural networks
The official Python client for the Huggingface Hub
Official inference library for Mistral models
Pytorch domain library for recommendation systems
20+ high-performance LLMs with recipes to pretrain, finetune at scale
A lightweight vision library for performing large object detection
Powering Amazon custom machine learning chips
Data manipulation and transformation for audio signal processing
Scripts for fine-tuning Meta Llama3 with composable FSDP & PEFT method
DoWhy is a Python library for causal inference
Multilingual Automatic Speech Recognition with word-level timestamps
Python Package for ML-Based Heterogeneous Treatment Effects Estimation
Probabilistic reasoning and statistical analysis in TensorFlow
PyTorch extensions for fast R&D prototyping and Kaggle farming