Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
State-of-the-art diffusion models for image and audio generation
Bring the notion of Model-as-a-Service to life
A library for accelerating Transformer models on NVIDIA GPUs
Library for OCR-related tasks powered by Deep Learning
Multilingual Automatic Speech Recognition with word-level timestamps
Superduper: Integrate AI models and machine learning workflows
A high-performance ML model serving framework, offers dynamic batching
Libraries for applying sparsification recipes to neural networks
State-of-the-art Parameter-Efficient Fine-Tuning
Optimizing inference proxy for LLMs
Neural Network Compression Framework for enhanced OpenVINO
Sparsity-aware deep learning inference runtime for CPUs
Large Language Model Text Generation Inference
Standardized Serverless ML Inference Platform on Kubernetes
Trainable models and NN optimization tools
A set of Docker images for training and serving models in TensorFlow
Efficient few-shot learning with Sentence Transformers
A toolkit to optimize ML models for deployment for Keras & TensorFlow
PyTorch extensions for fast R&D prototyping and Kaggle farming
Deep learning optimization library: makes distributed training easy
Official inference library for Mistral models
Open-source tool designed to enhance the efficiency of workloads
20+ high-performance LLMs with recipes to pretrain, finetune at scale
Data manipulation and transformation for audio signal processing