Port of Facebook's LLaMA model in C/C++
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU
Foundational Models for State-of-the-Art Speech and Text Translation
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine
Flux 2 image generation model pure C inference
Run the full 2.78-trillion-parameter Kimi K3 model
MiniMax H3 inference engine for Mac computers
FAIR Sequence Modeling Toolkit 2
Locally run an Instruction-Tuned Chat-Style LLM