Tokens / second / dollar
Build it. Buy it. Rent it.
Run any model three ways. Pick your model and how much you use it — we'll show you what's actually cheapest, scored on real tokens/sec and true cost.
~10M tokens is steady daily single-user use.
Your own GPU rig
$163/mo
18.7 tok/s · 2× NVIDIA GeForce RTX 4090
best case ~14s to first token on a 16k-token prompt
$143/mo hardware + $20 power · $5,130 upfront
Spec a buildA desktop AI box
$151/mo
15 tok/s · Mac Studio M3 Ultra (96GB)
best case ~2.6 min to first token on a 16k-token prompt
$147/mo hardware + $3 power · $5,299 upfront
tok/s estimated (bandwidth model)
Compare boxesA serverless cloud API
$3/mo
280 tok/s · Cloud API
TTFT unavailable (no published benchmark)
Pay per token — no hardware, scales to zero
per-token serverless API (output-token rate, nearest class)
Compare cloudFor Llama 3.1 70B at 10M/mo, rent is cheapest. Own a box for privacy, control, or always-on local access.
Single-stream estimates from third-party benchmarks. Build/Buy tok/s via our spec model; time to first token is a best-case estimate scaled from short-prompt llama.cpp prefill measurements (prefill slows at depth); rent is priced at per-token serverless output rates — input-token charges excluded on all paths.
The short answer: for most single-user workloads, renting a serverless API is cheapest — running a 70B model at ~10M tokens/month is roughly $3/month to rent versus a ~$3,500+ build or a ~$800+ desktop box amortized over years. Build or buy when you need privacy, always-on local access, or full control — your data never leaves your machine. Use the calculator to see the crossover for your model and volume.
100 parts · 13 GPUs · cloud + appliance + DIY · third-party tok/s
Get launch updates
Build guides, new GPU benchmarks, and price drops. No spam.
Three ways to run it. What each is actually like.
The calculator tells you what's cheapest. This is what you live with.
Build
Own the rig
A DIY GPU workstation. Total control, full privacy, and the lowest cost per token once it's busy — if you'll keep it busy and don't mind assembly, noise, and upkeep.
Spec a buildBrowse build guides →Buy
Plug it in
A turnkey desktop AI box — a Mac or appliance with big unified memory. Quiet, power-efficient, and it fits large models, for more money than parts alone.
Compare boxesRent
Just call the API
Serverless cloud inference billed per token. Nothing to own or maintain, scales to zero, fastest to start — cheapest for most personal use, but your data leaves your machine.
Compare cloudThe tradeoffs the price tag doesn't show.
| Build | Buy | Rent | |
|---|---|---|---|
| Setup | Assemble + configure | Unbox + sign in | An API key |
| Privacy | Fully local | Fully local | Data leaves your machine |
| Noise & power | Loud, 500W+ | Quiet, efficient | None (it's elsewhere) |
| Scales to zero | No — idle cost | No — idle cost | Yes — pay per token |
| Maintenance | You own it | Minimal | None |
| Biggest model | Add GPUs (VRAM) | Up to unified memory | Anything hosted |
| Starts at | ~$3.5k upfront | ~$800 upfront | $0 upfront |
Build vs Buy vs Rent — common questions
Is it cheaper to build, buy, or rent to run a local LLM?
- For most single-user use, renting a per-token cloud API is cheapest — amortized hardware for a build or a desktop box usually costs more per month than a serverless API at personal volumes. Owning (build or buy) wins on privacy, control, and always-on local access, and can pay off at sustained high utilization.
How much does it cost to run Llama 3.1 70B?
- At about 10M tokens per month, roughly $3/month on a serverless cloud API (output-token rate, cited 2026-07 class medians), versus about $158/month amortized for a 2× RTX 4090 build or about $61/month for a Mac Studio. Figures are single-stream estimates from third-party benchmarks.
When should I build a GPU rig instead of renting?
- Build when you need full data privacy, offline or always-on local access, control over the stack, or you run at sustained high volume where owning amortizes below per-token pricing. Renting wins on upfront cost ($0) and simplicity.
What's the difference between building, buying, and renting?
- Build is a DIY GPU workstation — the most control and the lowest cost per token when kept busy. Buy is a turnkey desktop AI box like a Mac — quiet, efficient, with big unified memory. Rent is a serverless cloud API billed per token — nothing to own, scales to zero, and the cheapest way to start.
We don't benchmark hardware. They do.
Every tok/s figure is resolved from third-party benchmarks and our grounded spec model — never our own numbers. See the methodology for how each estimate is built.
