DeepSeek V4.1-Flash Hits 494 Tokens/s on a 4x NVIDIA DGX Spark Home Rig
A home cluster of four NVIDIA DGX Spark units with ~512 GB of pooled memory ran the 552B-parameter DeepSeek V4.1-Flash model, reaching up to 494 tokens/s with 32 concurrent code requests. The rig costs between $25,000 and $40,000.
- Peak output of 494 tokens/s on code with 32 concurrent requests
- Single request: ~58 tokens/s for prose and ~96 tokens/s for code
- 552B-parameter model with 8B active parameters at prefill and 16B at decode
- Four 128 GB units bring the rig's total cost to $25,000-$40,000
Read next
Hardware