Fast inference engine for Transformer models
A high-performance ML model serving framework, offers dynamic batching
A GPU-accelerated library containing highly optimized building blocks
C++ library for high performance inference on NVIDIA GPUs
lightweight, standalone C++ inference engine for Google's Gemma models
MNN is a blazing fast, lightweight deep learning framework
Run frontier MoE models on hardware you already own
OpenVINO™ Toolkit repository
High-performance neural network inference framework for mobile
The Triton Inference Server provides an optimized cloud
High quality, fast, modular reference implementation of SSD in PyTorch
A computer vision framework to create and deploy apps in minutes
Fast and user-friendly runtime for transformer inference