Local Dev Starter
Single RTX 4080 SUPER build for running 7B–14B models locally with llama.cpp or Ollama.
~$2,800
~114.1 tok/s · 40.75 tok/s per $1k (Llama 3.1 8B)
View guideBuild templates
Curated AI workstation templates with tok/s and value ratings — customize any guide in the builder.
New here?
Answer 5 quick questions and get a personalized recommendation.
Single RTX 4080 SUPER build for running 7B–14B models locally with llama.cpp or Ollama.
~$2,800
~114.1 tok/s · 40.75 tok/s per $1k (Llama 3.1 8B)
View guideRTX 4090 powerhouse for 8B–34B models with headroom for agent workflows.
~$4,200
~93.6 tok/s · 22.29 tok/s per $1k (Qwen 2.5 14B)
View guideTwin 4090s pooling 48 GB VRAM to hold 70B-class models with real context headroom.
~$6,500
~18.7 tok/s · 2.88 tok/s per $1k (Llama 3.3 70B)
View guideTeam-grade dual 4090 rig targeting Llama 3.3 70B at Q4.
~$12,000
~14 tok/s · 1.17 tok/s per $1k (Llama 3.3 70B)
View guideDual-GPU workstation tuned for agent workloads with 16K context depth.
~$4,500
~337.5 tok/s · 75.00 tok/s per $1k (Qwen 3 Coder 30B-A3B (MoE))
View guideMinimal spend path to a solid Llama 3.1 8B daily driver.
~$1,800
~114.1 tok/s · 63.39 tok/s per $1k (Llama 3.1 8B)
View guideCost-conscious 14B inference box with modern single-GPU VRAM.
~$2,000
~65.2 tok/s · 32.60 tok/s per $1k (Qwen 2.5 14B)
View guide64GB RAM and RTX 4090 for LoRA fine-tuning on 8B–14B models.
~$4,800
~93.6 tok/s · 19.50 tok/s per $1k (Qwen 2.5 14B)
View guideHigh-memory build for concurrent team inference with vLLM on large models.
~$5,000
~14 tok/s · 2.80 tok/s per $1k (Llama 3.3 70B)
View guideNo assembly required
Prefer a ready-to-use device? These options ship configured and work out of the box.
Apple Silicon
16GB unified memory · ~28 tok/s on Llama 8B
View config →
Apple Silicon
48GB unified memory · ~42 tok/s on Llama 8B
View config →
Apple Silicon
36GB unified memory · ~41 tok/s on Llama 8B
View config →
Apple Silicon
96GB unified memory · ~68 tok/s on Llama 8B
View config →
NVIDIA Appliance
128GB unified memory · runs 70B at full precision
View on NVIDIA.com
Framework Appliance
128GB unified memory · runs 70B at full precision
View on Framework.com