Scripts for fine-tuning Meta Llama3 with composable FSDP & PEFT method
Unified Model Serving Framework
Multilingual Automatic Speech Recognition with word-level timestamps
Pytorch domain library for recommendation systems
PyTorch extensions for fast R&D prototyping and Kaggle farming
Create HTML profiling reports from pandas DataFrame objects
LMDeploy is a toolkit for compressing, deploying, and serving LLMs
A lightweight vision library for performing large object detection
Efficient few-shot learning with Sentence Transformers
Bring the notion of Model-as-a-Service to life
Phi-3.5 for Mac: Locally-run Vision and Language Models
Sparsity-aware deep learning inference runtime for CPUs
Superduper: Integrate AI models and machine learning workflows
A high-performance ML model serving framework, offers dynamic batching
Standardized Serverless ML Inference Platform on Kubernetes
MII makes low-latency and high-throughput inference possible
Deep learning optimization library: makes distributed training easy
Libraries for applying sparsification recipes to neural networks
Tensor search for humans
Library for serving Transformers models on Amazon SageMaker
A Unified Library for Parameter-Efficient Learning
Neural Network Compression Framework for enhanced OpenVINO
Easy-to-use Speech Toolkit including Self-Supervised Learning model
OffyAI — Local AI. Private by Design. Built for Your Machine.
GPU environment management and cluster orchestration