MiniMax H3 inference engine for Mac computers
TT-NN operator library, and TT-Metalium low level kernel programming
Run serverless GPU workloads with fast cold starts on bare-metal
DeepSeek 4 Flash local inference engine for Metal
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference
A scalable inference server for models optimized with OpenVINO
Fastest LLM inference runtime for Apple Silicon
gpt-oss-120b and gpt-oss-20b are two open-weight language models
LLM inference in C/C++
Run Bonsai (1-bit) and Ternary-Bonsai language models locally
QVAC Fabric: cross-platform LLM inference and fine-tuning
Nexa SDK is a comprehensive toolkit for supporting ONNX and GGML
Interface for OuteTTS models
Universal LLM Deployment Engine with ML Compilation
Run OpenClaw on a $5 chip
Training neural networks on Apple Neural Engine via APIs
MNN is a blazing fast, lightweight deep learning framework
LLM training in simple, raw C/CUDA
A high-performance inference engine for AI models
Open source solution that can meet the requirements of workloads
3D modeler, 3D game maker, 3D demo maker
Open platform for training, serving, and evaluating language models
An ecosystem of Rust libraries for working with large language models