Build guide
Single RTX 4080 SUPER build for running 7B–14B models locally with llama.cpp or Ollama.
Budget
$2,800
Profile
LOCAL DEV
Target model
Llama 3.1 8B
114.1tok/s
97–131.2 tok/s decode on Llama 3.1 8B
Value: 40.75 tok/s per $1k
Built for the solo developer who wants a reliable local Llama 3.1 8B or 14B model for code assist and daily chat — not production serving. The RTX 4080 SUPER's 16 GB VRAM fits both model families at Q4_K_M with room for a 4K KV cache, and it costs roughly $300 less than a 4090 while delivering over 90% of the throughput for 8B workloads. The Ryzen 7 7800X3D's 3D V-Cache cuts prefill latency noticeably at context windows above 2K, which matters more than raw clock speed for interactive use. Paired with 32 GB DDR5 and a 1 TB NVMe, the system handles model swaps without bottlenecks. What this build isn't: it won't run unquantized 34B models, it isn't designed for concurrent users, and LoRA fine-tuning on 14B will be slow. If those are requirements, step up to the Home Inference Workstation.