Port of Facebook's LLaMA model in C/C++
LLM inference in C/C++
Python bindings for llama.cpp
C#/.NET binding of llama.cpp, including LLaMa/GPT model inference
WebAssembly binding for llama.cpp - Enabling on-browser LLM inference
Run Local LLMs on Any Device. Open-source
Clippy, now with some AI
Run AI models locally on your machine with node.js bindings for llama
VS Code extension for LLM-assisted code/text completion
The free, Open Source alternative to OpenAI, Claude and others
Vim plugin for LLM-assisted code/text completion
Distribute and run LLMs with a single file
Open-source LLM load balancer and serving platform for hosting LLMs
Qwen3 is the large language model series developed by Qwen team
Performance-optimized AI inference on your GPUs
React and Electron-based app that executes the FreedomGPT LLM locally
Inference Llama 2 in one file of pure C
Run any Llama 2 locally with gradio UI on GPU or CPU from anywhere
Chinese LLaMA & Alpaca large language model + local CPU/GPU training
llama.go is like llama.cpp in pure Golang
Locally run an Instruction-Tuned Chat-Style LLM