Zhipu Opens GLM-5.3-FlashX Near 200 Tokens/s on ~100k Domestic Accelerators
Zhipu AI opened access to GLM-5.3-FlashX via API and testing center: peak generation about 200 tokens per second on approximately 100,000 Chinese AI accelerators. The base model GLM-5.3-Flash is a 320B-parameter MoE (18B active) with 1M token context, open-sourced on August 26.
- Peak generation speed of FlashX is about 200 tokens per second
- Acceleration provided by a cluster of ~100,000 Chinese AI accelerators
- Base GLM-5.3-Flash: 320B parameters, 18B active, 1M token context
- Stack optimized by Infra Agent on GLM-5.3: throughput increase about 3x
Read next
AI