Run models like Kimi-K2.5, GLM-5, DeepSeek, gpt-oss, Gemma, Qwen etc.
Port of Facebook's LLaMA model in C/C++
LLM inference in C/C++
Distribute and run LLMs with a single file
Run Local LLMs on Any Device. Open-source
The media player for language learning, with dual subtitles
Emscripten: An LLVM-to-WebAssembly Compiler
C#/.NET binding of llama.cpp, including LLaMa/GPT model inference
Production ready toolkit to run AI locally
Vector database plugin for Postgres, written in Rust
Next-gen AI+IoT framework for T2/T3/T5AI/ESP32/and more
Fast Multimodal LLM on Mobile Devices
AI-powered bridge connecting LLMs and advanced AI agents
Run a 1-billion parameter LLM on a $10 board with 256MB RAM
Integrate cutting-edge LLM technology quickly and easily into your app
Open-source large language model family from Tencent Hunyuan
TT-NN operator library, and TT-Metalium low level kernel programming
Inference Llama 2 in one file of pure C
High-speed Large Language Model Serving for Local Deployment
Mooncake is the serving platform for Kimi
INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model
Run PyTorch LLMs locally on servers, desktop and mobile
Research project. A Memory solution for users, teams, and applications
LLM training in simple, raw C/CUDA
The easiest way to use Ollama in .NET