Port of Facebook's LLaMA model in C/C++
Python bindings for llama.cpp
Run Local LLMs on Any Device. Open-source
Personal AI, On Personal Devices
Interface for OuteTTS models
Your Personal AI Assistant; easy to install, deploy on local or coud
Run a full local LLM stack with one command using Docker
Structured Outputs
Oobabooga - The definitive Web UI for local AI, with powerful features
Run Bonsai (1-bit) and Ternary-Bonsai language models locally
An easy-to-understand framework for LLM samplers
Qwen3 is the large language model series developed by Qwen team
Semantic ifs from open models, on a 3090 at home
First class Sublime Text AI assistant with gpt-5, Opus 4.6, Gemini 3
A proxy server for multiple ollama instances with Key security
Performance-optimized AI inference on your GPUs
GLM-4 series: Open Multilingual Multimodal Chat LMs
Towards Human-Sounding Speech
Run GGUF models easily with a UI or API. One File. Zero Install.
Spyder IDE plugin providing separate chat pane for AI Assistance
Inference Llama 2 in one file of pure C
FreeAskInternet is a completely free running search aggregator
Run any Llama 2 locally with gradio UI on GPU or CPU from anywhere
Chinese LLaMA & Alpaca large language model + local CPU/GPU training