Qwen3.8-Flash-Next vs Qwen3.8-27B: 125B Parameters, 6B Active — What the Qwen4 Preview Changes
Alibaba introduced the open MoE model Qwen3.8-Flash-Next: 125 billion parameters with 6 billion active per token and a 51 billion N-gram embedding table. Context is 262,144 tokens extendable to 1 million, claimed score of 62.5 on SWE-bench Pro, and training cost about 1/9 of Qwen3.7-Plus.
- 125B total / 6B active parameters, plus 51B N-gram embedding table
- Native context 262,144 tokens, extends to 1 million via YaRN
- 62.5 on SWE-bench Pro vs 61.7 for dense Qwen3.8-27B (vendor data)
- Claimed training cost about 1/9 of Qwen3.7-Plus
Read next
AI