Lightweight Python library for adding real-time multi-object tracking
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
Multilingual Automatic Speech Recognition with word-level timestamps
A set of Docker images for training and serving models in TensorFlow
Phi-3.5 for Mac: Locally-run Vision and Language Models
Superduper: Integrate AI models and machine learning workflows
A high-performance ML model serving framework, offers dynamic batching
State-of-the-art diffusion models for image and audio generation
Bring the notion of Model-as-a-Service to life
Libraries for applying sparsification recipes to neural networks
State-of-the-art Parameter-Efficient Fine-Tuning
Standardized Serverless ML Inference Platform on Kubernetes
Optimizing inference proxy for LLMs
LMDeploy is a toolkit for compressing, deploying, and serving LLMs
Neural Network Compression Framework for enhanced OpenVINO
Sparsity-aware deep learning inference runtime for CPUs
Large Language Model Text Generation Inference
Trainable models and NN optimization tools
Efficient few-shot learning with Sentence Transformers
Deep learning optimization library: makes distributed training easy
A toolkit to optimize ML models for deployment for Keras & TensorFlow
PyTorch extensions for fast R&D prototyping and Kaggle farming
Official inference library for Mistral models
Open-source tool designed to enhance the efficiency of workloads
MII makes low-latency and high-throughput inference possible