Port of OpenAI's Whisper model in C/C++
Fast inference engine for Transformer models
Low-latency AI inference engine optimized for mobile devices
A Python library for audio
Gradient boosting framework based on decision tree algorithms
High-speed Large Language Model Serving for Local Deployment
LLM inference in C/C++
LiteRT, successor to TensorFlow Lite
Bolt is a deep learning library with high performance
Easy-to-use deep learning framework with 3 key features
Lightweight inference library for ONNX files, written in C++
A High Performance Library for Sequence Processing and Generation
A RocksDB compatible KV storage engine with better performance
Fast and user-friendly runtime for transformer inference
10x faster matrix and vector operations