MiniMax H3 inference engine for Mac computers
TT-NN operator library, and TT-Metalium low level kernel programming
DeepSeek 4 Flash local inference engine for Metal
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
Fastest LLM inference runtime for Apple Silicon
Interface for OuteTTS models
Universal LLM Deployment Engine with ML Compilation
Run OpenClaw on a $5 chip
Training neural networks on Apple Neural Engine via APIs
A high-performance inference engine for AI models