jamesob's guide to running SOTA LLMs is a practical guide and configuration repository for running high-end language models on local hardware. It documents one developer’s local LLM setup, including hardware choices, GPU layout, storage, PCIe switches, kernel settings, and serving workflows. The repository compares budget levels ranging from dual RTX 3090 systems to high-end multi-GPU workstations with very large VRAM pools. It includes ready-to-run serving configurations for selected models...