MiniMax H3 inference engine for Mac computers
TT-NN operator library, and TT-Metalium low level kernel programming
DeepSeek 4 Flash local inference engine for Metal
Run serverless GPU workloads with fast cold starts on bare-metal
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
A scalable inference server for models optimized with OpenVINO
Fastest LLM inference runtime for Apple Silicon
gpt-oss-120b and gpt-oss-20b are two open-weight language models
Run Bonsai (1-bit) and Ternary-Bonsai language models locally
Nexa SDK is a comprehensive toolkit for supporting ONNX and GGML
QVAC Fabric: cross-platform LLM inference and fine-tuning
Interface for OuteTTS models
Universal LLM Deployment Engine with ML Compilation
Run OpenClaw on a $5 chip
Training neural networks on Apple Neural Engine via APIs
LLM training in simple, raw C/CUDA
A high-performance inference engine for AI models
Open source solution that can meet the requirements of workloads
3D modeler, 3D game maker, 3D demo maker
An ecosystem of Rust libraries for working with large language models