What should I buy?
Set your budget. We rank the hardware that gives you the most usable local AI for the money, and tell you honestly when the cloud is the smarter buy.
The honest part: for most people, a $200/mo cloud plan delivers more intelligence per dollar than any rig on this page. Buy local for privacy, control, offline use, or genuinely heavy daily generation, not to save money on day one.
9 builds under $6,000
Best value firstMac Studio M4 Max (128GB)
Best value128GB at 546 GB/s. The sweet spot for running 70B class models on one quiet box.
A $200/mo cloud sub gives you more tokens for less. Worth it for privacy and offline use, not cost.
MacBook Pro M4 Max (128GB)
Everything the Studio does, in a laptop. Pay a premium for portability.
A $200/mo cloud sub gives you more tokens for less. Worth it for privacy and offline use, not cost.
NVIDIA DGX Spark (128GB)
128GB of unified memory and CUDA, but 273 GB/s bandwidth means big models run slowly.
A $200/mo cloud sub gives you more tokens for less. Worth it for privacy and offline use, not cost.
RTX 5090 Build (32GB)
1792 GB/s of GDDR7, the fastest tokens per second you can buy in a single card.
Roughly trades even with a $200/mo cloud sub (~1.6 yr to break even). Buy it for control, not pure savings.
Mac Mini M4 Pro (64GB)
64GB of unified memory fits big models, but 273 GB/s caps how fast they generate.
A $200/mo cloud sub gives you more tokens for less. Worth it for privacy and offline use, not cost.
Dual RTX 3090 Rig (48GB)
48GB of fast VRAM by splitting across two used 3090s. The classic homelab inference rig.
A $200/mo cloud sub gives you more tokens for less. Worth it for privacy and offline use, not cost.
RTX 4090 Build (24GB)
Fastest 24GB single card. Great if your models fit in 24GB; you hit a wall above it.
A $200/mo cloud sub gives you more tokens for less. Worth it for privacy and offline use, not cost.
Used RTX 3090 Build (24GB)
The enthusiast value king. 936 GB/s of bandwidth for the price of a phone. CUDA + fast.
A $200/mo cloud sub gives you more tokens for less. Worth it for privacy and offline use, not cost.
Mac Mini M4 (24GB)
Cheapest real entry point. Silent, sips power, runs 7B to 14B models comfortably.
A $200/mo cloud sub gives you more tokens for less. Worth it for privacy and offline use, not cost.
Or rent instead
Not sure you'll run it enough to justify the box? Rent the same GPUs by the hour and only pay while you generate.
Prices and performance are estimates and change with the market. Some links are affiliate links; buying through them supports the tool at no extra cost to you. Token per second figures are approximations derived from memory bandwidth, not benchmarks.