What should I buy?

Set your budget. We rank the hardware that gives you the most usable local AI for the money, and tell you honestly when the cloud is the smarter buy.

$6,000
$

The honest part: for most people, a $200/mo cloud plan delivers more intelligence per dollar than any rig on this page. Buy local for privacy, control, offline use, or genuinely heavy daily generation, not to save money on day one.

9 builds under $6,000

Best value first

MacBook Pro M4 Max (128GB)

Apple Silicon
$4,699

Everything the Studio does, in a laptop. Pay a premium for portability.

Memory
128GB
Bandwidth
546GB/s
Power
90W
What it runs
Runs frontier 200B+ models
e.g. Qwen 3 Coder 30B-A3B (MoE) at 10.9 t/s · 141 models fit
vs cloudCloud wins on cost

A $200/mo cloud sub gives you more tokens for less. Worth it for privacy and offline use, not cost.

Buy on Apple · $4,699

NVIDIA DGX Spark (128GB)

DGX
$3,999

128GB of unified memory and CUDA, but 273 GB/s bandwidth means big models run slowly.

Memory
128GB
Bandwidth
273GB/s
Power
240W
What it runs
Runs frontier 200B+ models
e.g. Llama 3.1 8B at 10.2 t/s · 141 models fit
vs cloudCloud wins on cost

A $200/mo cloud sub gives you more tokens for less. Worth it for privacy and offline use, not cost.

Buy on NVIDIA · $3,999

RTX 5090 Build (32GB)

Single GPU
$3,500

1792 GB/s of GDDR7, the fastest tokens per second you can buy in a single card.

Memory
32GB
Bandwidth
1792GB/s
Power
650W
What it runs
Runs 70B class models
e.g. Gemma 4 26B-A4B (MoE) at 41.4 t/s · 127 models fit
vs cloudMarginal vs cloud

Roughly trades even with a $200/mo cloud sub (~1.6 yr to break even). Buy it for control, not pure savings.

Buy on Amazon · $3,500

Mac Mini M4 Pro (64GB)

Apple Silicon
$1,999

64GB of unified memory fits big models, but 273 GB/s caps how fast they generate.

Memory
64GB
Bandwidth
273GB/s
Power
45W
What it runs
Runs 100B+ models
e.g. Llama 3.1 8B at 10.2 t/s · 134 models fit
vs cloudCloud wins on cost

A $200/mo cloud sub gives you more tokens for less. Worth it for privacy and offline use, not cost.

Buy on Apple · $1,999

Dual RTX 3090 Rig (48GB)

Multi-GPU
$3,800

48GB of fast VRAM by splitting across two used 3090s. The classic homelab inference rig.

Memory
48GB
Bandwidth
936GB/s
Power
800W
What it runs
Runs 100B+ models
e.g. Falcon 40B at 14.0 t/s · 129 models fit
vs cloudCloud wins on cost

A $200/mo cloud sub gives you more tokens for less. Worth it for privacy and offline use, not cost.

Buy on Amazon · $3,800

RTX 4090 Build (24GB)

Single GPU
$2,800

Fastest 24GB single card. Great if your models fit in 24GB; you hit a wall above it.

Memory
24GB
Bandwidth
1008GB/s
Power
500W
What it runs
Runs 30B class models
e.g. InternLM 2.5 20B at 30.2 t/s · 116 models fit
vs cloudCloud wins on cost

A $200/mo cloud sub gives you more tokens for less. Worth it for privacy and offline use, not cost.

Buy on Amazon · $2,800

Used RTX 3090 Build (24GB)

Single GPU
$1,700

The enthusiast value king. 936 GB/s of bandwidth for the price of a phone. CUDA + fast.

Memory
24GB
Bandwidth
936GB/s
Power
400W
What it runs
Runs 30B class models
e.g. InternLM 2.5 20B at 28.1 t/s · 116 models fit
vs cloudCloud wins on cost

A $200/mo cloud sub gives you more tokens for less. Worth it for privacy and offline use, not cost.

Buy on Amazon · $1,700

Mac Mini M4 (24GB)

Apple Silicon
$999

Cheapest real entry point. Silent, sips power, runs 7B to 14B models comfortably.

Memory
24GB
Bandwidth
120GB/s
Power
35W
What it runs
Runs 30B class models
e.g. Qwen 2.5 3B at 12.0 t/s · 116 models fit
vs cloudCloud wins on cost

A $200/mo cloud sub gives you more tokens for less. Worth it for privacy and offline use, not cost.

Buy on Apple · $999

Or rent instead

Not sure you'll run it enough to justify the box? Rent the same GPUs by the hour and only pay while you generate.

Prices and performance are estimates and change with the market. Some links are affiliate links; buying through them supports the tool at no extra cost to you. Token per second figures are approximations derived from memory bandwidth, not benchmarks.