Kimi K3 in C
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU
...Memory presets balance pinned layers and an expert LRU cache for laptops, desktops, workstations, and servers. Incremental generation preserves KV cache and recurrent state between tokens. The repository also includes tokenization, safetensors loading, diagnostics, benchmarks, trace replay, and comparisons against a PyTorch reference.