Neural Network Compression Framework for enhanced OpenVINO
Large Language Model Text Generation Inference
PyTorch extensions for fast R&D prototyping and Kaggle farming
LMDeploy is a toolkit for compressing, deploying, and serving LLMs
20+ high-performance LLMs with recipes to pretrain, finetune at scale
Efficient few-shot learning with Sentence Transformers
AIMET is a library that provides advanced quantization and compression
Multilingual Automatic Speech Recognition with word-level timestamps
Python Package for ML-Based Heterogeneous Treatment Effects Estimation
Pytorch domain library for recommendation systems
Open-source tool designed to enhance the efficiency of workloads
Tensor search for humans
Data manipulation and transformation for audio signal processing
A Pythonic framework to simplify AI service building
Adversarial Robustness Toolbox (ART) - Python Library for ML security
A Unified Library for Parameter-Efficient Learning
Operating LLMs in production
Replace OpenAI GPT with another LLM in your app
Training and deploying machine learning models on Amazon SageMaker
Uplift modeling and causal inference with machine learning algorithms
State-of-the-art Parameter-Efficient Fine-Tuning
The Triton Inference Server provides an optimized cloud
Optimizing inference proxy for LLMs
MII makes low-latency and high-throughput inference possible
A set of Docker images for training and serving models in TensorFlow