MiniMax H3 inference engine for Mac computers
TT-NN operator library, and TT-Metalium low level kernel programming
DeepSeek 4 Flash local inference engine for Metal
Run serverless GPU workloads with fast cold starts on bare-metal
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference
Fastest LLM inference runtime for Apple Silicon
LLM inference in C/C++
gpt-oss-120b and gpt-oss-20b are two open-weight language models
Run Bonsai (1-bit) and Ternary-Bonsai language models locally
Nexa SDK is a comprehensive toolkit for supporting ONNX and GGML
QVAC Fabric: cross-platform LLM inference and fine-tuning
Interface for OuteTTS models
Universal LLM Deployment Engine with ML Compilation
Run OpenClaw on a $5 chip
Training neural networks on Apple Neural Engine via APIs
LLM training in simple, raw C/CUDA
A high-performance inference engine for AI models
Open source solution that can meet the requirements of workloads
3D modeler, 3D game maker, 3D demo maker
Open platform for training, serving, and evaluating language models
An ecosystem of Rust libraries for working with large language models
Cheat sheet for Google Cloud developers