The Triton Inference Server provides an optimized cloud
A high-performance ML model serving framework, offers dynamic batching
Bring the notion of Model-as-a-Service to life
Open platform for training, serving, and evaluating language models
OpenMMLab Model Deployment Framework
Run any Llama 2 locally with gradio UI on GPU or CPU from anywhere
Lightweight anchor-free object detection model
Deploy a ML inference service on a budget in 10 lines of code