A 2.78-trillion-parameter Kimi K3 running inference on a single CPU
Port of OpenAI's Whisper model in C/C++
Running large language models on a single GPU
Supercharge Your LLM with the Fastest KV Cache Layer
Find the local LLM that actually runs and performs best
A system monitoring tool that exposes system metrics
Chat, serve, monitor, and connect MLX models from one macOS app
Run the full 2.78-trillion-parameter Kimi K3 model
Open deep learning compiler stack for cpu, gpu