Build guide
Minimal spend path to a solid Llama 3.1 8B daily driver.
Budget
$1,800
Profile
LOCAL DEV
Target model
Llama 3.1 8B
114.1tok/s
97–131.2 tok/s decode on Llama 3.1 8B
Value: 63.39 tok/s per $1k
If you want a fast, reliable Llama 3.1 8B setup and nothing more, this is the minimum sensible build. Llama 3.1 8B at Q4_K_M needs about 6 GB of VRAM, so any modern discrete GPU with 8+ GB works — but the RTX 4080 SUPER gives substantial headroom and upgrade room to 14B later without touching anything else. You'll see 80–100+ tok/s on 8B at Q4 with this card, which is faster than most people can read. The 7800X3D is arguably overkill for pure 8B workloads, but the price delta over a budget CPU is small at this tier and future-proofs the platform. 32 GB RAM is comfortable for Ollama or llama.cpp plus a browser and IDE running simultaneously. This build is intentionally a daily driver, not a research platform — it doesn't have the VRAM headroom for 34B or 70B models regardless of quantization, and it won't serve multiple concurrent users.