Port of Facebook's LLaMA model in C/C++
LLM inference in C/C++
Python bindings for llama.cpp
C#/.NET binding of llama.cpp, including LLaMa/GPT model inference
WebAssembly binding for llama.cpp - Enabling on-browser LLM inference
Run Local LLMs on Any Device. Open-source
Personal AI, On Personal Devices
Clippy, now with some AI
Run AI models locally on your machine with node.js bindings for llama
VS Code extension for LLM-assisted code/text completion
Interface for OuteTTS models
The free, Open Source alternative to OpenAI, Claude and others
Vim plugin for LLM-assisted code/text completion
Run a full local LLM stack with one command using Docker
Your Personal AI Assistant; easy to install, deploy on local or coud
Terminal-native coding agent powered by local LLMs
Distribute and run LLMs with a single file
QVAC Fabric: cross-platform LLM inference and fine-tuning
Open-source LLM load balancer and serving platform for hosting LLMs
Run Bonsai (1-bit) and Ternary-Bonsai language models locally
DevoxxGenie is a plugin for IntelliJ IDEA that uses local LLM's
An easy-to-understand framework for LLM samplers
Oobabooga - The definitive Web UI for local AI, with powerful features
Qwen3 is the large language model series developed by Qwen team