Run a 1-billion parameter LLM on a $10 board with 256MB RAM
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
157 models, 30 providers, one command to find what runs on hardware
Real-time NVIDIA GPU dashboard
Explore large language models in 512MB of RAM
Generate embeddings from large-scale graph-structured data
Compact 3B-param multimodal model for efficient on-device reasoning
Lightweight 24B agentic coding model with vision and long context