Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
Run the full 2.78-trillion-parameter Kimi K3 model
llama.go is like llama.cpp in pure Golang
Locally run an Instruction-Tuned Chat-Style LLM
Generate embeddings from large-scale graph-structured data
Compact 3B-param multimodal model for efficient on-device reasoning
Lightweight 24B agentic coding model with vision and long context