Build guide
RTX 4090 powerhouse for 8B–34B models with headroom for agent workflows.
Budget
$4,200
Profile
HOME INFERENCE
Target model
Qwen 2.5 14B
93.6tok/s
79.6–107.6 tok/s decode on Qwen 2.5 14B
Value: 22.29 tok/s per $1k
Designed for the home power user who regularly switches between 8B, 14B, and 34B models and runs multi-step agent pipelines where latency compounds. The RTX 4090's 24 GB VRAM is the single-card ceiling for consumer hardware — it runs Qwen 2.5 14B at full Q8 precision, or a quantized 34B with short context, without any VRAM pressure. At Q4_K_M on 8B models, the 4090 benchmarks at 60–80 tok/s, making real-time agentic loops feel snappy. The Ryzen 7 7800X3D handles prefill spikes from long agent prompts without becoming the bottleneck. This build assumes one user running one model at a time. It isn't built for serving concurrent requests from multiple users — vLLM's continuous batching needs more VRAM headroom than a single 4090 provides for 34B models. For teams, see the Team Server or Dual 4090 builds.