Phi-3.5 for Mac: Locally-run Vision and Language Models
State-of-the-art diffusion models for image and audio generation
Superduper: Integrate AI models and machine learning workflows
A high-performance ML model serving framework, offers dynamic batching
Libraries for applying sparsification recipes to neural networks
Fast inference engine for Transformer models
Standardized Serverless ML Inference Platform on Kubernetes
State-of-the-art Parameter-Efficient Fine-Tuning
Optimizing inference proxy for LLMs
LMDeploy is a toolkit for compressing, deploying, and serving LLMs
Neural Network Compression Framework for enhanced OpenVINO
Sparsity-aware deep learning inference runtime for CPUs
Large Language Model Text Generation Inference
A set of Docker images for training and serving models in TensorFlow
Trainable models and NN optimization tools
Efficient few-shot learning with Sentence Transformers
Multilingual Automatic Speech Recognition with word-level timestamps
Uncover insights, surface problems, monitor, and fine tune your LLM
A toolkit to optimize ML models for deployment for Keras & TensorFlow
PyTorch extensions for fast R&D prototyping and Kaggle farming
The Triton Inference Server provides an optimized cloud
Official inference library for Mistral models
Deep learning optimization library: makes distributed training easy
MII makes low-latency and high-throughput inference possible
20+ high-performance LLMs with recipes to pretrain, finetune at scale