chiprook
← AI
AIOctober 4, 2026, 23:58

$1,600 RTX 4090 hits 100 T/s on 125B Qwen model, beating H100

A volunteer team ran the 125B-parameter Qwen 3.8 Flash Next model on a single RTX 4090, reaching roughly 100 T tokens/s aggregate throughput. The setup uses int4 quantization, speculative decoding with a 4B drafter and CUDA-graph-fused inference, cutting cost per 1M tokens to $0.00004.

$1,600 RTX 4090 hits 100 T/s on 125B Qwen model, beating H100
#Nvidia#RTX4090#Qwen#H100
Read next
AI

Strata runs a 125B LLM on a gaming PC with 12 GB VRAM

AI

Qwen3.8-Flash-Next vs Qwen3.8-27B: 125B Parameters, 6B Active — What the Qwen4 Preview Changes

Gadgets

Razer Reportedly Quotes $3,695 to Repair a 3-Year-Old RTX 4090 Laptop

Hardware

RTX 4090 12VHPWR Adapter Melts Despite 80% Power Limit