Can I run it?
Pick a model. We'll show you the cheapest hardware that runs it, the box we'd actually recommend, and the fastest option, with real token per second estimates.
GLM 4.6 355B-A32B (MoE)•355B•GLM
GeneralCodingReasoningMathCreative
Runs on 7 builds
RecommendedFastest
NVIDIA DGX Station (GB300, 784GB)
DGX$80,000
Q4_K_M•medium
210.0 GB784 GB
Good•24.3 t/s
Minimum
Mac Studio M3 Ultra (256GB)
Apple Silicon$7,499
Q4_K_M•medium
210.0 GB256 GB
Very Slow•2.8 t/s
Every build that runs GLM 4.6 355B-A32B (MoE)
Mac Studio M3 Ultra (256GB)
Apple Silicon · Q4_K_M
$7,499
Mac Studio M3 Ultra (512GB)
Apple Silicon · Q4_K_M
$9,499
4x DGX Spark Cluster (512GB)
DGX · Q4_K_M
$16,000
2x Mac Studio M3 Ultra (1TB cluster)
Apple Silicon · Q4_K_M
$19,000
8x RTX PRO 6000 Blackwell (768GB)
Multi GPU · Q4_K_M
$78,000
NVIDIA DGX Station (GB300, 784GB)
DGX · Q4_K_M
$80,000
8x NVIDIA H200 SXM (1.1TB)
Multi GPU · Q4_K_M
$300,000
Token per second figures are estimates derived from memory bandwidth, not benchmarks. "Recommended" means the cheapest build that runs this model at good quality and usable speed. Some links are affiliate links.