A Pythonic framework to simplify AI service building
Integrate, train and manage any AI models and APIs with your database
Lightweight Python library for adding real-time multi-object tracking
MII makes low-latency and high-throughput inference possible
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
Scripts for fine-tuning Meta Llama3 with composable FSDP & PEFT method
Bring the notion of Model-as-a-Service to life
Superduper: Integrate AI models and machine learning workflows
A high-performance ML model serving framework, offers dynamic batching
AIMET is a library that provides advanced quantization and compression
Optimizing inference proxy for LLMs
Neural Network Compression Framework for enhanced OpenVINO
Large Language Model Text Generation Inference
Probabilistic reasoning and statistical analysis in TensorFlow
A lightweight vision library for performing large object detection
Easiest and laziest way for building multi-agent LLMs applications
Efficient few-shot learning with Sentence Transformers
Multilingual Automatic Speech Recognition with word-level timestamps
Uncover insights, surface problems, monitor, and fine tune your LLM
The Triton Inference Server provides an optimized cloud
Open-source tool designed to enhance the efficiency of workloads
Standardized Serverless ML Inference Platform on Kubernetes
Data manipulation and transformation for audio signal processing
Phi-3.5 for Mac: Locally-run Vision and Language Models
A Unified Library for Parameter-Efficient Learning