Run a 1-billion parameter LLM on a $10 board with 256MB RAM
157 models, 30 providers, one command to find what runs on hardware
Find the local LLM that actually runs and performs best
Real-time NVIDIA GPU dashboard
Explore large language models in 512MB of RAM
llama.go is like llama.cpp in pure Golang
Locally run an Instruction-Tuned Chat-Style LLM