Run Local LLMs on Any Device. Open-source
ONNX Runtime: cross-platform, high performance ML inferencing
Protect and discover secrets using Gitleaks
C++ library for high performance inference on NVIDIA GPUs
High-performance neural network inference framework for mobile
Official inference library for Mistral models
An MLOps framework to package, deploy, monitor and manage models
MNN is a blazing fast, lightweight deep learning framework
Standardized Serverless ML Inference Platform on Kubernetes
A GPU-accelerated library containing highly optimized building blocks
Neural Network Compression Framework for enhanced OpenVINO
Easy-to-use deep learning framework with 3 key features
A set of Docker images for training and serving models in TensorFlow
Powering Amazon custom machine learning chips
Library for OCR-related tasks powered by Deep Learning
Set of comprehensive computer vision & machine intelligence libraries
AIMET is a library that provides advanced quantization and compression
A general-purpose probabilistic programming system
Library for serving Transformers models on Amazon SageMaker
Superduper: Integrate AI models and machine learning workflows
A unified framework for scalable computing
Unified Model Serving Framework
The Triton Inference Server provides an optimized cloud
Build Production-ready Agentic Workflow with Natural Language
LLM.swift is a simple and readable library