Towards Human-Sounding Speech
Performance-optimized AI inference on your GPUs
GLM-4 series: Open Multilingual Multimodal Chat LMs
Lightweight, standalone, multi-platform, and privacy focused local LLM
React and Electron-based app that executes the FreedomGPT LLM locally
Inference Llama 2 in one file of pure C
Run GGUF models easily with a UI or API. One File. Zero Install.
Amica is an open source interface for interactive communication
Run any Llama 2 locally with gradio UI on GPU or CPU from anywhere
Open source large-language-model based code completion engine
Self-hosted ChatGPT-like chatbot powered by Llama models locally
Chinese LLaMA & Alpaca large language model + local CPU/GPU training
Chat with your favourite LLaMA models in a native macOS app
llama.go is like llama.cpp in pure Golang
Locally run an Instruction-Tuned Chat-Style LLM
Fast uncensored Gemma model optimized for local chat and coding
JetBrains’ 4B parameter code model for completions
Jan-v1-edge: efficient 1.7B reasoning model optimized for edge devices