Operating LLMs in production
Integrate, train and manage any AI models and APIs with your database
PyTorch extensions for fast R&D prototyping and Kaggle farming
Sparsity-aware deep learning inference runtime for CPUs
A lightweight vision library for performing large object detection
A toolkit to optimize ML models for deployment for Keras & TensorFlow
Trainable, memory-efficient, and GPU-friendly PyTorch reproduction
Easiest and laziest way for building multi-agent LLMs applications
Trainable models and NN optimization tools
Official inference library for Mistral models
Low-latency REST API for serving text-embeddings
State-of-the-art Parameter-Efficient Fine-Tuning
Data manipulation and transformation for audio signal processing
Libraries for applying sparsification recipes to neural networks
Multi-Modal Neural Networks for Semantic Search, based on Mid-Fusion
Multilingual Automatic Speech Recognition with word-level timestamps
Uncover insights, surface problems, monitor, and fine tune your LLM
Unified Model Serving Framework
Optimizing inference proxy for LLMs
Superduper: Integrate AI models and machine learning workflows
Efficient few-shot learning with Sentence Transformers
A library for accelerating Transformer models on NVIDIA GPUs
Simplifies the local serving of AI models from any source
Replace OpenAI GPT with another LLM in your app
State-of-the-art diffusion models for image and audio generation