Port of Facebook's LLaMA model in C/C++
Python bindings for llama.cpp
C#/.NET binding of llama.cpp, including LLaMa/GPT model inference
Run Bonsai (1-bit) and Ternary-Bonsai language models locally
Qwen3 is the large language model series developed by Qwen team
Chinese LLaMA & Alpaca large language model + local CPU/GPU training
llama.go is like llama.cpp in pure Golang
Locally run an Instruction-Tuned Chat-Style LLM
Fast uncensored Gemma model optimized for local chat and coding
JetBrains’ 4B parameter code model for completions
Jan-v1-edge: efficient 1.7B reasoning model optimized for edge devices