Skip to content
ClankerBuilder
Sign in

Tokens / second / dollar

Build it. Buy it. Rent it.

Run any model three ways. Pick your model and how much you use it — we'll show you what's actually cheapest, scored on real tokens/sec and true cost.

~10M tokens is steady daily single-user use.

Build

Your own GPU rig

$163/mo

18.7 tok/s · 2× NVIDIA GeForce RTX 4090

best case ~14s to first token on a 16k-token prompt

$143/mo hardware + $20 power · $5,130 upfront

Spec a build
Buy

A desktop AI box

$151/mo

15 tok/s · Mac Studio M3 Ultra (96GB)

best case ~2.6 min to first token on a 16k-token prompt

$147/mo hardware + $3 power · $5,299 upfront

tok/s estimated (bandwidth model)

Compare boxes
RentCheapest

A serverless cloud API

$3/mo

280 tok/s · Cloud API

TTFT unavailable (no published benchmark)

Pay per token — no hardware, scales to zero

per-token serverless API (output-token rate, nearest class)

Compare cloud

For Llama 3.1 70B at 10M/mo, rent is cheapest. Own a box for privacy, control, or always-on local access.

Single-stream estimates from third-party benchmarks. Build/Buy tok/s via our spec model; time to first token is a best-case estimate scaled from short-prompt llama.cpp prefill measurements (prefill slows at depth); rent is priced at per-token serverless output rates — input-token charges excluded on all paths.

The short answer: for most single-user workloads, renting a serverless API is cheapest — running a 70B model at ~10M tokens/month is roughly $3/month to rent versus a ~$3,500+ build or a ~$800+ desktop box amortized over years. Build or buy when you need privacy, always-on local access, or full control — your data never leaves your machine. Use the calculator to see the crossover for your model and volume.

100 parts · 13 GPUs · cloud + appliance + DIY · third-party tok/s

Get launch updates

Build guides, new GPU benchmarks, and price drops. No spam.

Three ways to run it. What each is actually like.

The calculator tells you what's cheapest. This is what you live with.

Build

Own the rig

A DIY GPU workstation. Total control, full privacy, and the lowest cost per token once it's busy — if you'll keep it busy and don't mind assembly, noise, and upkeep.

Spec a buildBrowse build guides →

Buy

Plug it in

A turnkey desktop AI box — a Mac or appliance with big unified memory. Quiet, power-efficient, and it fits large models, for more money than parts alone.

Compare boxes

Rent

Just call the API

Serverless cloud inference billed per token. Nothing to own or maintain, scales to zero, fastest to start — cheapest for most personal use, but your data leaves your machine.

Compare cloud

The tradeoffs the price tag doesn't show.

Build vs Buy vs Rent: operational tradeoffs
BuildBuyRent
SetupAssemble + configureUnbox + sign inAn API key
PrivacyFully localFully localData leaves your machine
Noise & powerLoud, 500W+Quiet, efficientNone (it's elsewhere)
Scales to zeroNo — idle costNo — idle costYes — pay per token
MaintenanceYou own itMinimalNone
Biggest modelAdd GPUs (VRAM)Up to unified memoryAnything hosted
Starts at~$3.5k upfront~$800 upfront$0 upfront

Build vs Buy vs Rent — common questions

Is it cheaper to build, buy, or rent to run a local LLM?

For most single-user use, renting a per-token cloud API is cheapest — amortized hardware for a build or a desktop box usually costs more per month than a serverless API at personal volumes. Owning (build or buy) wins on privacy, control, and always-on local access, and can pay off at sustained high utilization.

How much does it cost to run Llama 3.1 70B?

At about 10M tokens per month, roughly $3/month on a serverless cloud API (output-token rate, cited 2026-07 class medians), versus about $158/month amortized for a 2× RTX 4090 build or about $61/month for a Mac Studio. Figures are single-stream estimates from third-party benchmarks.

When should I build a GPU rig instead of renting?

Build when you need full data privacy, offline or always-on local access, control over the stack, or you run at sustained high volume where owning amortizes below per-token pricing. Renting wins on upfront cost ($0) and simplicity.

What's the difference between building, buying, and renting?

Build is a DIY GPU workstation — the most control and the lowest cost per token when kept busy. Buy is a turnkey desktop AI box like a Mac — quiet, efficient, with big unified memory. Rent is a serverless cloud API billed per token — nothing to own, scales to zero, and the cheapest way to start.

We don't benchmark hardware. They do.

Every tok/s figure is resolved from third-party benchmarks and our grounded spec model — never our own numbers. See the methodology for how each estimate is built.

LocalScorer/LocalLLaMAllama.cppTom's HardwareServeTheHomeArtificial AnalysisTechPowerUpNotebookCheck