Flux 2 image generation model pure C inference
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine
Run the full 2.78-trillion-parameter Kimi K3 model
Fast, Sharp & Reliable Agentic Intelligence
FlashMLA: Efficient Multi-head Latent Attention Kernels
Real-time behaviour synthesis with MuJoCo, using Predictive Control
ChatGPT integration with Unity Editor