$1,600 RTX 4090 hits 100 T/s on 125B Qwen model, beating H100
A volunteer team ran the 125B-parameter Qwen 3.8 Flash Next model on a single RTX 4090, reaching roughly 100 T tokens/s aggregate throughput. The setup uses int4 quantization, speculative decoding with a 4B drafter and CUDA-graph-fused inference, cutting cost per 1M tokens to $0.00004.
- int4 quantization cut VRAM use from 30+ GB to about 22 GB
- Speculative decoding with a 4B drafter added roughly 20% speed
- Cost per 1M tokens is $0.00004 versus $0.00009 on H100
- Energy efficiency is about 285 Gtokens/kWh versus 240 on H100
Read next
AI