Port of OpenAI's Whisper model in C/C++
Serve, optimize and scale PyTorch models in production
Simplifies the local serving of AI models from any source
A scalable inference server for models optimized with OpenVINO
Low-latency REST API for serving text-embeddings
Unified Model Serving Framework
The Triton Inference Server provides an optimized cloud
Deep Learning API and Server in C++14 support for Caffe, PyTorch
Openai style api for open large language models
Deploy a ML inference service on a budget in 10 lines of code