A toolkit to optimize ML models for deployment for Keras & TensorFlow
Replace OpenAI GPT with another LLM in your app
Sparsity-aware deep learning inference runtime for CPUs
The Triton Inference Server provides an optimized cloud
Trainable, memory-efficient, and GPU-friendly PyTorch reproduction
A Pythonic framework to simplify AI service building
Easy-to-use Speech Toolkit including Self-Supervised Learning model
Tensor search for humans
A unified framework for scalable computing
Standardized Serverless ML Inference Platform on Kubernetes
AIMET is a library that provides advanced quantization and compression
Images to inference with no labeling
Open platform for training, serving, and evaluating language models
A computer vision framework to create and deploy apps in minutes