Ongoing research training transformer models at scale
ONNX Runtime: cross-platform, high performance ML inferencing
PArallel Distributed Deep LEarning: Machine Learning Framework
High-performance neural network inference framework for mobile
Trainable models and NN optimization tools
OpenVINO™ Toolkit repository
A GPU-accelerated library containing highly optimized building blocks
Probabilistic reasoning and statistical analysis in TensorFlow
MII makes low-latency and high-throughput inference possible
Deep Learning API and Server in C++14 support for Caffe, PyTorch
A unified framework for scalable computing
Connect MATLAB to LLM APIs, including OpenAI® Chat Completions
Library for serving Transformers models on Amazon SageMaker
Powering Amazon custom machine learning chips
A set of Docker images for training and serving models in TensorFlow
Deep learning optimization library: makes distributed training easy
Easy-to-use deep learning framework with 3 key features
OpenMMLab Model Deployment Framework
RNN with great LLM performance
A computer vision framework to create and deploy apps in minutes
Implementation of model parallel autoregressive transformers on GPUs
The deep learning toolkit for speech-to-text
Guide to deploying deep-learning inference networks
Toolkit for allowing inference and serving with MXNet in SageMaker